top | item 46752212

(no title)

How well does such llm research hold up as new models are released?

discuss

dexdal|1 month ago

Most model research decays because the evaluation harness isn’t treated as a stable artefact. If you freeze the tasks, acceptance criteria, and measurement method, you can swap models and still compare apples to apples. Without that, each release forces a reset and people mistake novelty for progress.