Evaluating AI Agents: 7 Mistakes That Make Your Tests Lie
Published: 2026-09-25 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - The point: a green status on eval tests doesn't guarantee an AI agent works correctly; checking only the final answer misses actions that were never performed. - Where to find it: an article by Sergey Proshchaev published on Habr in the OTUS company blog. - Limitation: eval tests create a false sense of reliability; real evaluation requires checking system state and intermediate steps. ### 🔍 What was found In the article "Evaluating AI Agents: 7 Mistakes That Make Your Tests Lie," Sergey Proshchaev, Tech Lead of the Java|Kotlin track in FinTech & E-commerce and an OTUS instructor, breaks down typical failures in eval harnesses. He gives a telling example: a team has an eval set of 40 tasks, the CI run holds around 90% success, but two weeks later support brings in a chat where the agent reported "request submitted" while the request database is empty. The tests aren't broken — they're simply measuring the wrong thing. The first mistake is counting success by the final text rather than the system state. The model can perfectly describe a result that never happened. In τ-bench from Sierra, the main criterion is whether the final database state matches the target, and in the original experiment GPT-4o in this configuration solved less than half the tasks. The second mistake is a test set made up entirely of happy paths; production throws hostile scenarios at you: prompt injections, jailbreaks, stale data, structurally valid but semantically wrong tool responses. The author separately highlights retries: a timeout is not the same as failure, and with a non-idempotent tool a diligent agent can charge a card twice. ### 💡 Why it matters Evaluating an agent is not a single number but several orthogonal measurements. A correct final state doesn't mean a correct process: an agent may cancel an order but fail to check permissions. Practical takeaways: verify the state transition, not just its existence, cross-check logs and intermediate steps, and add at least one broken scenario to the eval set for every happy-path task. That's the only way to see whether the agent actually does what's needed. ### 🧩 Context A few months ago the author already covered six architectural mistakes that keep agents from surviving to launch — that article was about how an agent is built. This one covers the adjacent layer: how to understand whether an agent actually works. At the end, Sergey promises a summary table, a checklist, and a conclusion about which skill the rake garden is really testing.
⚡ The gist in 5 seconds - The point: a green status on eval tests doesn't guarantee an AI agent works correctly; checking only the final answer misses actions that were never performed.
- Where to find it: an article by Sergey Proshchaev published on Habr in the OTUS company blog.
- Limitation: eval tests create a false sense of reliability; real evaluation requires checking system state and intermediate steps.
🔍 What was found In the article "Evaluating AI Agents: 7 Mistakes That Make Your Tests Lie," Sergey Proshchaev, Tech Lead of the Java|Kotlin track in FinTech & E-commerce and an OTUS instructor, breaks down typical failures in eval harnesses.
He gives a telling example: a team has an eval set of 40 tasks, the CI run holds around 90% success, but two weeks later support brings in a chat where the agent reported "request submitted" while the request database is empty.
The tests aren't broken — they're simply measuring the wrong thing.
The first mistake is counting success by the final text rather than the system state.
The model can perfectly describe a result that never happened.