Back to work

GUIDEEvaluation / Release gates

How do we evaluate consequential AI decisions before every release?

How to turn real failure modes into a release suite covering routes, evidence, tool use, specialist judgement and complete user outcomes.

FOR
Research, engineering, safety and product teams
FORMAT
Practical implementation guide
OUTCOME
Turn approved real-world failures and uncertain cases into new tests while monitoring that the suite remains representative.
01

UNDERSTAND THE PROBLEM

Begin with the failure you need to prevent.

A general benchmark can tell you whether a model is capable. It cannot tell you whether your product chose the correct route for your user, evidence and intended use.

Consequential AI needs product-specific evaluation. The suite should represent the decisions the system makes, the hazards it must catch, the tools it may use and the cases where the correct behaviour is to ask, abstain or escalate.

The unit of evaluation is therefore not always one answer. It may be a multi-turn trajectory ending in a real product outcome.

02

DESIGN THE SYSTEM

Make the operating rules explicit.

A dependable AI system is easier to build when the team can see the decisions, evidence, boundaries and ownership around it. The following principles turn an ambiguous ambition into components that can be implemented and reviewed.

Build evaluations from failures the business cares about.

Begin with costly wrong routes, missed hazards, unsupported actions and operational incidents rather than a generic list of language tasks.

Grade components and complete journeys

Test retrieval, routing and tool calls separately for diagnosis, then test the end-to-end outcome users experience.

Use deterministic graders first

Check schemas, routes, citations, permissions and prohibited actions mechanically. Use model or specialist graders for the judgement that remains.

Measure different failures separately

A missed hazard and an unnecessary escalation have different consequences. One average quality score hides both.

03

IMPLEMENT IN ORDER

Build the smallest complete loop.

Do not automate every adjacent task at once. Start with one valuable journey, carry it from signal to outcome, and preserve enough evidence to know whether it worked. Expand only after that loop is dependable.

  1. 01

    Write the failure catalogue

    Gather incidents, expert concerns, support cases and red-team examples, then connect each to a product or business consequence.

  2. 02

    Create representative cases

    Include ordinary language, ambiguity, long context, adversarial phrasing, tool failure and examples that should not complete automatically.

  3. 03

    Define graders and thresholds

    State what can be checked exactly, what needs a rubric and which critical failure blocks a release regardless of the average score.

  4. 04

    Run every relevant change

    Trigger the suite for model, prompt, knowledge, rule, tool and orchestration releases, not only model upgrades.

  5. 05

    Feed production back

    Turn approved real-world failures and uncertain cases into new tests while monitoring that the suite remains representative.

04

KNOW WHEN IT WORKS

Measure behaviour, not how impressive the demo looks.

The useful measure is whether the system creates the intended business or product outcome while staying inside its boundary. Review these checks before launch and whenever the model, data, prompt, tools or workflow changes.

Critical recall

The system catches the high-impact routes where missing one case is unacceptable.

Over-escalation

Safety controls do not make the product unusable by routing every ordinary case to a person.

Trajectory quality

The path, tool use and final state are correct, not merely the last sentence.

Release discipline

Critical regressions block deployment and remain visible to the accountable team.

The work behind the work.

  1. Anthropic: Demystifying evals for AI agents
  2. OpenAI: HealthBench
  3. NIST AI Risk Management Framework

Related questions

CASE STUDY

How do we build AI products that can operate in regulated domains?

What does your AI product need to do reliably?

Bring us the difficult product, model or workflow problem. We bring the senior product, design and engineering team required to get it working in production.

Start a conversation