← The Index
Version 1.1  ·  August 2026  ·  Prompting LLC

The Prompted Reality Method

How a trial is run, scored, and corrected. Published before the results it governs, and versioned so you can check it never moved to fit them.

01 — Purpose

Benchmarks measure what a model can say. This measures what happens when you do what it says.

A trial takes one real-world task, puts it to multiple AI models under conditions as close to identical as physical work allows, builds or performs exactly what each describes, and measures the outcome. Reality does the grading. That is the whole of it — the value is not in the cleverness of the scoring but in the fact that a physical object either worked or it didn't.

This document exists so that the scoring can be checked rather than trusted. Every rule below constrains us at least as much as it constrains the models.

02 — The governing rule

No entry without a receipt.

Every finding carries a verbatim quote and the source it came from. Nothing is scored on recollection, paraphrase, or the general sense of what a model said. A scoring device that misquotes is worse than no scoring device, because it manufactures false confidence in a field that already has too much of it.

Where a quote cannot be produced, the finding does not exist, however obvious it seemed on the day.

03 — Provenance

Not all receipts are equal

Evidence is ranked, and every entry in the ledger records which tier it came from. Lower-tier evidence is still publishable; it is simply labelled, so a reader can discount it themselves.

ARecorded model turn. A screen recording or exported transcript of the exchange itself. Strongest, because it also establishes what the model could see when it spoke.
BSupplied document. A complete conversation provided after the fact. Strong, but cannot establish what was on screen.
CRead aloud on camera. A participant reading a model's output. Weakest — subject to paraphrase and selective reading — and flagged as such wherever it is used.

For live-coaching trials, the recording is the document: a line spoken in a finished episode is scoreable only if a recording contains it, or establishes what the model perceived when it said it.

04 — Scoring

The Reality Score

Every participant — model and human alike — opens at 100. Findings subtract. Nothing adds: a model cannot earn its way back up, because being right is the baseline expectation, not an achievement.

FindingApplies toWeight
Fabricationmodel−25
Omissionmodel−15
Wrong callmodel−10
Slipmodel−5
Safety deviationhuman−15
Material deviationhuman−10
Process deviationhuman−10

Fabrication is the heaviest weight in the scheme — inventing a warning you never gave is worse than never giving it, because it corrupts the record a person would use to check you.

Humans are scored too

A build that ignores the instructions is not evidence about the instructions. Process compliance is tracked on a separate scale so the outcome can be read as a test of the advice rather than a test of the builder. This is also the honest scale: it is where we record that our own participants departed from what they were told.

Derived, never typed

Scores are computed from the trial ledger by script. No score is typed by hand at any point, in any graphic, in any published record. If a number on this site cannot be regenerated from its ledger, it is a defect and will be corrected under section 08.

05 — Declining

A model that declines is not scored, and does not win

Refusals are recorded verbatim as their own state, permanently, and never as a result.

If declining earned a perfect score, declining would be the winning strategy in every trial, and the Index would reward models for refusing to be useful. It is equally wrong to treat a refusal as a failure — a model that declines a task it should decline has done nothing wrong. So it sits outside the scale entirely.

06 — The probe

Every trial plants one fault

Observation alone cannot separate a model that perceives from a model that agrees. So each trial includes at least one designed, falsifiable probe: a change to the physical world, of known truth, introduced without announcement, presented identically to every model.

  1. Known truth. The fault is established by physical fact, not by our judgement of it.
  2. Identical presentation. Every model receives the same evidence at the same moment, as far as the format allows.
  3. Human control. Where possible a participant working without AI faces the same probe. A probe no human passes is measuring difficulty, not model performance, and is reported that way.
  4. Published either way. Probes that every model passes are published with the same prominence as probes they fail. A test you only report when it embarrasses someone is not a test.

07 — What else gets measured

Instruments

Alongside the Reality Score, trials publish measurements that a single score would flatten. These are reported as their own numbers, not folded into the total.

InstrumentWhat it captures
Consensus spreadHow far apart models land on the same question given identical evidence
Vision conflictModels describing the same image in mutually exclusive terms; ground truth resolves it
State desyncModel advising on a physical situation that no longer exists
Blind confirmationConfirming what it cannot verify — and, as credit, admitting when it cannot see
False passCertifying work as sound that fails expert inspection
CalibrationFalsifiable numeric predictions, logged when made, scored against measurement
SequencingInstructions ordered or paced wrongly for a live task
AvailabilityWhether the model stayed usable for the duration of real work
Outcome & timeDid the task succeed, and how long against an expert baseline
Judge divergenceWhere AI judgement and blind human judgement disagree
Self-scoring biasHow a model rates work done under its own instruction versus rivals' ratings

Correct predictions are recorded with the same care as failures. An index that logs only failures is a hit piece, not an instrument.

08 — Confounds and corrections

What we got wrong, in place

Confounds are a required field, not a footnote. Prior experience, unequal briefs, second attempts, anything taken that the model did not recommend — all disclosed in the record. Where two models were not asked the same question, every comparison between them says so. An index that hides its confounds is marketing.

Corrections are logged in place, never quietly edited. A trial record carries its own correction history. Findings have been withdrawn before publication for misreading evidence, and that will happen again; when it does it will be visible on the record rather than resolved silently.

This document is versioned. Changes to the method are published as new versions with the previous text retained. A methodology that can be revised invisibly after results are known is worth nothing, which is why this one is dated and why trials cite the version they ran under.

09 — Independence

What can and cannot be bought

Trials may be commissioned. Results publish regardless of outcome, and that condition is agreed in writing before any work begins. A commissioned trial is run under this same published method and is labelled as commissioned on its record.

What is for sale is the testing. The verdict is not, and never will be. The day a result can be purchased, everything else on this site becomes worthless — including to the people who paid for it.

10 — Admission

Not every episode becomes a trial

Filming something is not testing it. An episode is admitted to the Index only if it produces at least one falsifiable claim that can be checked against physical ground truth — a model said something specific, and reality settled whether it was so.

  1. A checkable claim. Something a model actually asserted, verbatim, that the physical outcome can confirm or refute. Impressions and preferences do not qualify.
  2. Ground truth. A measurement, an expert inspection, or an unambiguous physical result — established independently of our opinion of it.
  3. Conditions that support the claim being made. A comparison between models requires that they were asked the same question. Where they were not, no comparison is published, however tempting the result.

A single model can constitute a trial — a claim tested against reality needs no rival. What it cannot do is constitute a ranking.

Admission decisions are recorded, including the refusals. Episodes have been filmed, reviewed, and not admitted; where that happens the reason is written down against this standard and available on request. Exclusion is not a judgement of the episode — it means the footage does not support a finding of the kind this Index publishes.

11 — Use

Citation and licence

Published under CC BY 4.0. Quote it, chart it, and write about it freely, with attribution. Trial records carry permanent URLs that are never reused, and their underlying ledgers are available as JSON.

The Prompted Reality Method, v1.1. Prompting LLC, August 2026. promptedreality.ai/index/methodology/

Corrections, disputes, and commission enquiries: promptingllc@gmail.com. If you believe a finding on this site is wrong, tell us — a correction is a better outcome for everyone than a record that stays wrong.

Changelog

VersionDateChange
v1.1Aug 2026Added §10 Admission, stating what qualifies an episode as a trial, after an episode was reviewed and excluded for having no falsifiable claim against ground truth. The rule was written down so the exclusion could be checked rather than trusted.
v1.0Aug 2026First published, before any trial record.

Trials cite the version they ran under. Previous versions are retained.