Skip to main content

Service

Evaluation & assurance

Measurement designed before the system exists, and the evidence pack that lets your second line say yes.

Most AI systems are measured after they are built, against whatever they happen to be good at. That is how organisations end up with impressive demonstrations and disappointing production.

We design the measurement first: an adjudicated dataset built with the people who do the work today, a rubric more than one reviewer can apply, agreement measured and reported, and a regression gate that blocks changes which reduce quality.

What you get

Named artefacts, not a slide pack. Each one is a thing your team can open, run, or hand to an auditor.

ArtefactWhat it contains
Golden datasetAdjudicated cases with documented reasoning, built with your subject-matter experts. Routine, known-hard, and adversarial cases in deliberate proportion.
Rubric + agreement studyScoring criteria plus measured inter-rater agreement, which bounds every claim anyone can make afterwards.
Offline suiteOutcome and trajectory scoring, tool-call precision and recall, cost and latency per resolved task.
Judge calibration reportWhere a model judge is used: its agreement with human review, its failure modes, and the re-calibration schedule.
Regression gateWired into CI so a change that reduces measured quality cannot merge.
Adversarial suitePrompt injection, tool abuse, and boundary-violation cases maintained as a standing test set.
Evidence packThe documentation set a second-line function needs: method, data provenance, results with uncertainty, limitations, and oversight design.

How it runs

  1. 01 · 1–2 weeks

    Design

    What counts as success, who adjudicates, and what would falsify the claim.

  2. 02 · 2–4 weeks

    Build the set

    Case collection, adjudication, agreement measurement, documentation.

  3. 03 · 2–4 weeks

    Instrument

    Harness, CI integration, production sampling and review workflow.

  4. 04 · 1–2 weeks

    Evidence

    The written pack, and a session with the function that has to accept it.

What we use

Chosen per engagement against your constraints. We have no reseller relationships and no incentive to recommend one of these over another.

Evaluation

Custom harnessesTrajectory scoringModel-as-judgeStatistical testingHuman review workflows

Delivery

GitHub ActionsGitLab CIArgo Workflows

Governance frameworks

ISO/IEC 42001NIST AI RMFEU AI Act

What we do not do

Knowing where our usefulness stops saves everyone a procurement cycle.

  • We do not report a score without the agreement measurement behind it. A two-point improvement inside a ten-point noise band is not a result.
  • We do not use the same model as judge and subject. The correlation in failure modes is real and it flatters.
  • We do not sign off systems. We produce evidence; your risk function decides.

Questions we are asked

Is model-as-judge trustworthy?

Within limits that are not optional. Calibrate against human judgement on a held-out sample before trusting it, report the agreement rate alongside any result it produced, never use the same model as judge and subject, and keep a permanently human-reviewed subset because judge drift is silent. Used that way it lets you evaluate thousands of cases weekly instead of forty.

How large does the evaluation set need to be?

Large enough that the difference you care about is detectable given your reviewers’ agreement rate, which is a calculation, not a rule of thumb. In practice a few hundred well-adjudicated cases outperform several thousand casually labelled ones.

Can you evaluate a system you did not build?

Yes, and it is a common engagement. We usually start by reading the existing suite and writing up honestly what it can and cannot support.

Related

Start with the constraint.

Most of these projects are shaped by what you cannot do rather than what you want. Data that cannot leave the estate, a model you cannot host with a third party, a decision somebody has to justify to a regulator. Tell us yours and we will say honestly whether we can work inside it.