← Prompted Reality

The Prompted Reality Index

A permanent record of what AI models say to do, and what happens when someone actually does it.

What this is

Benchmarks measure what a model can say. The Index measures what happens when you do what it says.

Each trial puts the same real-world problem to multiple AI models, builds what they describe, and measures the outcome (velocity, distance, success, failure) under a published, versioned methodology. Every model quote is verbatim from a recording or supplied document. Every finding carries its receipt.

The rule that governs this Index: no entry without a receipt. A scoring device that misquotes is worse than no scoring device. Scores are derived from the trial ledger by script — never typed by hand. A model that declines to answer is recorded as its own state: never scored, and never counted as a win.

Trials

Records are sealed until their episode releases, then published here permanently at a fixed URL under CC BY 4.0. Numbers are assigned at publication, in the order records appear on this Index, so an unpublished trial has no number to reuse.

Commission a trial

Have an AI product, agent, or claim you want tested in the real world?

We run commissioned trials under the same published methodology as everything else here — and results publish regardless of outcome. That condition is not negotiable; it is the product. promptingllc@gmail.com