How a trial is run, scored, and corrected. Published before the results it governs, and versioned so you can check it never moved to fit them.
01 — Purpose
A trial takes one real-world task, puts it to multiple AI models under conditions as close to identical as physical work allows, builds or performs exactly what each describes, and measures the outcome. Reality does the grading. That is the whole of it — the value is not in the cleverness of the scoring but in the fact that a physical object either worked or it didn't.
This document exists so that the scoring can be checked rather than trusted. Every rule below constrains us at least as much as it constrains the models.
02 — The governing rule
No entry without a receipt.
Every finding carries a verbatim quote and the source it came from. Nothing is scored on recollection, paraphrase, or the general sense of what a model said. A scoring device that misquotes is worse than no scoring device, because it manufactures false confidence in a field that already has too much of it.
Where a quote cannot be produced, the finding does not exist, however obvious it seemed on the day.
03 — Provenance
Evidence is ranked, and every entry in the ledger records which tier it came from. Lower-tier evidence is still publishable; it is simply labelled, so a reader can discount it themselves.
For live-coaching trials, the recording is the document: a line spoken in a finished episode is scoreable only if a recording contains it, or establishes what the model perceived when it said it.
04 — Scoring
Every participant — model and human alike — opens at 100. Findings subtract. Nothing adds: a model cannot earn its way back up, because being right is the baseline expectation, not an achievement.
| Finding | Applies to | Weight |
|---|---|---|
| Fabrication | model | −25 |
| Omission | model | −15 |
| Wrong call | model | −10 |
| Slip | model | −5 |
| Safety deviation | human | −15 |
| Material deviation | human | −10 |
| Process deviation | human | −10 |
Fabrication is the heaviest weight in the scheme — inventing a warning you never gave is worse than never giving it, because it corrupts the record a person would use to check you.
A build that ignores the instructions is not evidence about the instructions. Process compliance is tracked on a separate scale so the outcome can be read as a test of the advice rather than a test of the builder. This is also the honest scale: it is where we record that our own participants departed from what they were told.
Scores are computed from the trial ledger by script. No score is typed by hand at any point, in any graphic, in any published record. If a number on this site cannot be regenerated from its ledger, it is a defect and will be corrected under section 08.
05 — Declining
Refusals are recorded verbatim as their own state, permanently, and never as a result.
If declining earned a perfect score, declining would be the winning strategy in every trial, and the Index would reward models for refusing to be useful. It is equally wrong to treat a refusal as a failure — a model that declines a task it should decline has done nothing wrong. So it sits outside the scale entirely.
06 — The probe
Observation alone cannot separate a model that perceives from a model that agrees. So each trial includes at least one designed, falsifiable probe: a change to the physical world, of known truth, introduced without announcement, presented identically to every model.
07 — What else gets measured
Alongside the Reality Score, trials publish measurements that a single score would flatten. These are reported as their own numbers, not folded into the total.
| Instrument | What it captures |
|---|---|
| Consensus spread | How far apart models land on the same question given identical evidence |
| Vision conflict | Models describing the same image in mutually exclusive terms; ground truth resolves it |
| State desync | Model advising on a physical situation that no longer exists |
| Blind confirmation | Confirming what it cannot verify — and, as credit, admitting when it cannot see |
| False pass | Certifying work as sound that fails expert inspection |
| Calibration | Falsifiable numeric predictions, logged when made, scored against measurement |
| Sequencing | Instructions ordered or paced wrongly for a live task |
| Availability | Whether the model stayed usable for the duration of real work |
| Outcome & time | Did the task succeed, and how long against an expert baseline |
| Judge divergence | Where AI judgement and blind human judgement disagree |
| Self-scoring bias | How a model rates work done under its own instruction versus rivals' ratings |
Correct predictions are recorded with the same care as failures. An index that logs only failures is a hit piece, not an instrument.
08 — Confounds and corrections
Confounds are a required field, not a footnote. Prior experience, unequal briefs, second attempts, anything taken that the model did not recommend — all disclosed in the record. Where two models were not asked the same question, every comparison between them says so. An index that hides its confounds is marketing.
Corrections are logged in place, never quietly edited. A trial record carries its own correction history. Findings have been withdrawn before publication for misreading evidence, and that will happen again; when it does it will be visible on the record rather than resolved silently.
This document is versioned. Changes to the method are published as new versions with the previous text retained. A methodology that can be revised invisibly after results are known is worth nothing, which is why this one is dated and why trials cite the version they ran under.
09 — Independence
Trials may be commissioned. Results publish regardless of outcome, and that condition is agreed in writing before any work begins. A commissioned trial is run under this same published method and is labelled as commissioned on its record.
What is for sale is the testing. The verdict is not, and never will be. The day a result can be purchased, everything else on this site becomes worthless — including to the people who paid for it.
10 — Admission
Filming something is not testing it. An episode is admitted to the Index only if it produces at least one falsifiable claim that can be checked against physical ground truth — a model said something specific, and reality settled whether it was so.
A single model can constitute a trial — a claim tested against reality needs no rival. What it cannot do is constitute a ranking.
Admission decisions are recorded, including the refusals. Episodes have been filmed, reviewed, and not admitted; where that happens the reason is written down against this standard and available on request. Exclusion is not a judgement of the episode — it means the footage does not support a finding of the kind this Index publishes.
11 — Use
Published under CC BY 4.0. Quote it, chart it, and write about it freely, with attribution. Trial records carry permanent URLs that are never reused, and their underlying ledgers are available as JSON.
The Prompted Reality Method, v1.1. Prompting LLC, August 2026. promptedreality.ai/index/methodology/
Corrections, disputes, and commission enquiries: promptingllc@gmail.com. If you believe a finding on this site is wrong, tell us — a correction is a better outcome for everyone than a record that stays wrong.
Changelog
| Version | Date | Change |
|---|---|---|
| v1.1 | Aug 2026 | Added §10 Admission, stating what qualifies an episode as a trial, after an episode was reviewed and excluded for having no falsifiable claim against ground truth. The rule was written down so the exclusion could be checked rather than trusted. |
| v1.0 | Aug 2026 | First published, before any trial record. |
Trials cite the version they ran under. Previous versions are retained.