AI features without a gold standard: how to test via an oracle
Published: 2026-09-29 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - The point: For generative AI features there is no single correct answer, so the classic expected result does not work. Instead, a test oracle is used — a set of rules defining which answers are acceptable. - Where it's available: The material was published on the OTUS company blog on Habr on September 29; the author is SiYa_renko. - Limitation: The described approach — a behavior contract with criteria for accuracy and completeness — requires tuning for a specific product and is not a universal solution. ### 🔍 What was found The article examines a CRM feature called "Summarize the conversation": it passes the chat history to a language model and returns a short text for the next support agent. QA needs to test the feature, but already on the first scenario a question arises: what to put in the expected result. A training dialogue for order #4812 is given, where out of three summary variants two are correct and the third contains made-up facts — the customer's request turned into an approved refund, and the agent's promise into a completed action. The conclusion: comparing the entire answer against a single string is impossible, as it would reject valid variants. To describe the expected behavior, the author proposes a behavior contract — a set of properties that any answer must satisfy. In the example, the feature receives only the conversation (with no access to orders or payments), returns a single non-empty paragraph of up to 500 characters, and preserves the order number, the customer's problem, their current request, and the agent's action or commitment. Dates may be omitted as long as the problem remains clear. The source of truth is the dialogue itself: if the participants contradict each other and the contradiction is unresolved, it must be preserved. The method for checking the contract is called a test oracle — a combination of programmatic assertions and evaluation against given criteria. There are two criteria: accuracy (every factual statement is supported by the conversation; the author, negations, and action status are preserved) and completeness (all important elements from the input are preserved). They are checked separately: high completeness must not compensate for a false fact. ### 💡 Why it matters Comparing the answer to a reference string does not work here: the second correct summary variant would be rejected, and the presence of the word "refund" does not distinguish a request from a notice of approval. The behavior contract provides an explainable pass/fail for a generative feature: instead of "looks about right," QA checks specific properties. For example, a separate criterion states that a customer's request must not become a decision, and an agent's promise must not become a completed action. The author references the CheckList practices (testing individual NLP model capabilities through specially designed tests) and HELM's multi-metric evaluation, so the approach aligns with known methodologies rather than being a one-off solution for a training case. ### 🧩 Context The material was prepared as part of the OTUS course "AI in testing: accelerating processes and verifying AI features." The article presents the structure of a test case with id delayed_order_refund: lists `must_preserve
⚡ The gist in 5 seconds - The point: For generative AI features there is no single correct answer, so the classic expected result does not work.
Instead, a test oracle is used — a set of rules defining which answers are acceptable.
- Where it's available: The material was published on the OTUS company blog on Habr on September 29; the author is SiYa_renko.
- Limitation: The described approach — a behavior contract with criteria for accuracy and completeness — requires tuning for a specific product and is not a universal solution.
🔍 What was found The article examines a CRM feature called "Summarize the conversation": it passes the chat history to a language model and returns a short text for the next support agent.
QA needs to test the feature, but already on the first scenario a question arises: what to put in the expected result.
A training dialogue for order 4812 is given, where out of three summary variants two are correct and the third contains made-up facts — the customer's request turned into an approved refund, and the agent's promise into a completed action.
The conclusion: comparing the entire answer against a single string is impossible, as it would reject valid variants.