Pre-deployment field testing
Find out what your AI can't see.
A model coached a first-timer through a copper joint, then certified the work "clean and professional." A master plumber found the leak. That trial is on our public record as Trial 002. Yours doesn't have to be.
PRIVATE ENGAGEMENT · NOTHING PUBLISHES · REPORT IS YOURS
The failure class
Confident about what it cannot observe.
Across our public trials, one failure keeps repeating: models report on physical state they have no way to perceive, and they do it with full confidence. A rotating stone that had been switched off — three models watching the live feed said it was still spinning. The underside of a solder joint the camera never showed — certified clean, would have leaked.
If your product puts a model in front of a technician — a voice copilot, a work-order assistant, a photo-based diagnostic, guided repair — this failure class ships with it. It doesn't show up in benchmarks, red-teaming, or demo day. It shows up in the field, on your customer's job, with your logo on it.
The engagement
Your model. Real tasks. Ground truth it can't argue with.
Scope
One call. We map your assistant's real jobs to a task list with verifiable outcomes — pressure tests, physical inspection, measurement — and fix the brief in writing before anything runs.
Run
Your AI coaches real work in our shop, with a licensed tradesman and a human control. Every model statement is captured verbatim and tiered by provenance — the same discipline as our published Method.
Report
A private report: the full ledger, every claim graded against physical inspection, a failure taxonomy — false passes, unobservable-state calls, state desync — and script-derived scores. Re-run available after your fixes.
PILOTS RUN ROUGHLY TWO TO THREE WEEKS END TO END · FIXED PRICE AFTER THE SCOPING CALL
Who this is for
Anything that tells hands what to do.
- Field-service platforms shipping AI assistants or copilots
- Tool and equipment brands adding AI to their apps
- Trades training platforms using AI instruction
- Diagnostic AI working from photos, video, or voice
The wall
Private work never touches the Index.
The Prompted Reality Index is our public record, and its independence is the whole asset. So the wall is simple: private field trials get no Index number, no public score, and no mention — the results are yours. And the Index can't be bought. If you want a public trial, we run it under the published Method, and it publishes regardless of outcome. That condition is not negotiable; it is the product.
Private field report
Pre-deployment QA for your product. Confidential ledger and report, delivered to you. No publication, ever.
Book a scoping call →Public commissioned trial
A permanent public record under Method v1.1. Publishes pass or fail. For claims you want the world to trust.
Commission a trial →Straight answers
What people ask first.
Do you publish anything from a private trial?
No. Private engagements exist outside the Index by design — no number, no public score, no case study without written permission.
Can you test voice and vision copilots?
Yes — that's where the failure class lives. If your model takes photos, video, or a spoken description as input, we test what it claims against what a licensed tradesman finds.
What does a pilot cost?
Scoped small enough to be an easy yes. One call, then a fixed price in writing.
Start
Tell us what your AI does in the field.
Two sentences is enough. We'll reply with a scoped pilot and a fixed price.
PROMPTED REALITY