Founding pilot now open

Your demo passed.
The holdout may not.

We pressure-test one important AI workflow before customers find the expensive failures. You get evidence your team can reproduce and a straight answer on what to do next.

One workflowDeliberately narrow
Held-out casesKept out of tuning
Traceable failuresInputs and evidence preserved
A clear decisionShip, narrow or stop
FIX THE TEST BEFORE THE PRODUCTKEEP THE HOLDOUT UNTOUCHEDSHOW THE ACTUAL FAILURELET THE ANSWER BE STOP

A hard test of one claim your product makes.

Not a broad consultancy exercise. Not a compliance badge. A compact independent test that gives a small team something useful enough to change a product decision.

01

A test set built around the real job

We turn one live product claim into realistic cases, awkward edge cases and explicit pass or fail rules.

02

An untouched holdout

The final cases stay out of tuning. That shows whether a fix generalises or merely learns the examples it has already seen.

03

Failures your team can reproduce

Every material omission, invention, wrong action or overconfident answer comes with the input and evidence needed to inspect it.

04

A decision, not another dashboard

You get the fixes that matter first and a plain view on whether to ship, narrow the workflow or stop.

The headline score looked excellent. The product still failed.

In our own synthetic technical experiment, a conversational resolver detected nearly every food component and matched every structured source. It looked ready to improve.

The untouched holdout exposed systematic undercounting and a dangerous false-reassurance problem. We stopped the build. Strong averages had hidden the failure that mattered.

This was an internal synthetic engineering test, not a customer engagement, clinical validation or independent certification. It is shown because we think stopped work is evidence too.
FINAL HOLDOUTSTOP
95.7%component precision and recall
100%structured source resolution
−7.79%systematic calorie bias
26.67%dangerous false reassurance

The attractive metrics did not cancel the material failure.

Small enough to finish. Serious enough to change the answer.

01

Pick one promise

We agree the single workflow, who relies on it and what a costly failure looks like.

02

Set the rules

Cases and thresholds are fixed before the final run, including a holdout neither side tunes against.

03

Try to break it

We run realistic and adversarial cases, preserve the evidence and separate product failures from test noise.

04

Make the call

You receive the scorecard, reproducible failures and the shortest sensible repair order.

One owner. One workflow. One costly way to be wrong.

01 Your AI feature already works well enough to test, not just describe.

02 Its output or action can be checked against evidence or explicit rules.

03 A miss, invention or wrong action has a real product or operating cost.

04 You want to find the limit, not purchase a flattering score.

Tell us the claim you cannot afford to get wrong.

Send three things: what the AI does, who relies on it and what a costly failure looks like. We will tell you whether a compact independent test can produce a useful decision before asking you to buy one.

£950 UK
US$1,250 International

[email protected] The first conversation is a fit check, not a sales performance. If the scope is weak, we will say so.