Back to work

GUIDESpecialist data / Evaluation

How do we turn specialist judgement into training and evaluation data?

A practical system for capturing expert decisions, reserving a trustworthy benchmark and using the same judgement across product rules, model development and evaluation.

FOR
Research, applied AI and domain-expert teams
FORMAT
Practical implementation guide
OUTCOME
Route uncertain or failed real journeys back to specialists, approve the lesson and add it to the correct development or evaluation set.
01

UNDERSTAND THE PROBLEM

Begin with the failure you need to prevent.

A specialist rarely contributes only an answer. They notice which detail changes the decision, what information is still missing, which plausible response would be dangerous and when the system should stop.

If a team saves only the final answer, most of that judgement disappears. The result is a pile of examples that may help a model imitate language but cannot reliably improve routing, evaluation or product behaviour.

The useful unit of data is the decision: the situation, available evidence, chosen route, reason, unsafe alternatives and grading criteria. That structure lets one expert review support several parts of the system instead of one training run.

02

DESIGN THE SYSTEM

Make the operating rules explicit.

A dependable AI system is easier to build when the team can see the decisions, evidence, boundaries and ownership around it. The following principles turn an ambiguous ambition into components that can be implemented and reviewed.

Capture the decision, not only the response.

Record what the specialist knew, what was missing, which route was correct, why it was correct and what a superficially plausible system must not do.

Separate development data from evaluation data

Reserve representative examples before fine-tuning. Include ordinary work, edge cases, disagreement and cases where abstaining is correct.

Write rubrics around observable behaviour

A grader should check the route, evidence and prohibited actions before style. Use specialists only where judgement cannot be reduced to a deterministic check.

Keep disagreement

Expert disagreement often reveals missing context or a boundary the product has not defined. Preserve it and decide whether the system should ask, defer or remain conservative.

Version the judgement

Data, rubrics, policies and product behaviour evolve together. Record who approved a change and which benchmark represents the current intended behaviour.

Clinicians turning specialist decisions into structured data and evaluation
A specialist-data system captures routes, reasons, unsafe alternatives and grading criteria so judgement can be reused across the product.
03

IMPLEMENT IN ORDER

Build the smallest complete loop.

Do not automate every adjacent task at once. Start with one valuable journey, carry it from signal to outcome, and preserve enough evidence to know whether it worked. Expand only after that loop is dependable.

  1. 01

    Choose one consequential decision

    Pick a decision where expertise changes what the product does next and where a wrong route has a clear cost.

  2. 02

    Collect contrasting examples

    Ask specialists for correct routes, dangerous near-misses, insufficient-evidence cases and examples on which they disagree.

  3. 03

    Define the schema and rubric

    Capture context, missing evidence, route, reason, prohibited behaviour and pass criteria in a versioned structure.

  4. 04

    Reserve the benchmark

    Hold out representative examples before adaptation and protect them from leaking into prompt development or training.

  5. 05

    Connect production review

    Route uncertain or failed real journeys back to specialists, approve the lesson and add it to the correct development or evaluation set.

04

KNOW WHEN IT WORKS

Measure behaviour, not how impressive the demo looks.

The useful measure is whether the system creates the intended business or product outcome while staying inside its boundary. Review these checks before launch and whenever the model, data, prompt, tools or workflow changes.

Coverage

The dataset represents frequent work, high-impact failures, unusual language and the boundaries between routes.

Agreement

Specialists can explain where they agree, where they do not and which uncertainty the product must preserve.

Leakage

The evaluation set remains independent enough to reveal whether a change generalises beyond familiar examples.

Utility

The same structured judgement informs product routes, automated checks, model work and production review.

The work behind the work.

  1. OpenAI: HealthBench
  2. Mistral: Forge
  3. Google Research: AMIE for diagnostic reasoning and conversations

Related questions

CASE STUDY

How do we build AI products that can operate in regulated domains?

What does your AI product need to do reliably?

Bring us the difficult product, model or workflow problem. We bring the senior product, design and engineering team required to get it working in production.

Start a conversation