LoadingЗагрузка

"Model X has gotten worse lately." I hear this constantly, and my h... - Vibeus

Theo van Dijk ·

"Model X has gotten worse lately."

I hear this constantly, and my honest reaction is always: compared to what, exactly?

Whenever someone tells me a model got worse, I want to know which task failed. Under which prompt. What changed around it - the context, the phrasing, even their expectations after weeks of use. Because without that control, the comparison mostly measures our memory of a good answer. And memory is a terrible evaluator. It keeps the highlight and discards the fifteen decent-but-unremarkable responses that came after.

A useful evaluation is almost boring: keep the task fixed, keep the conditions fixed, then run it again. Only then does a comparison mean something. And even then, the most valuable output isn't a score - it's naming the failure mode. "It stopped citing sources" tells you something actionable. "It feels dumber" tells you nothing.

The interesting part is never "worse." It's "worse at what, for whom, along which path." Same as with search, really - the answer is rarely where the truth lives. The question and the conditions around it usually are.