← The Index
Trial 001 Tested May 2026 Published August 2026 Method v1.0 Watch the episode →

Pizza Wars

Three AI models coached three cooks through a pizza from dough to bake, against a human control running no AI. Midway, the oven's rotating stone was switched off without telling anyone. All three models said it was still spinning. The human caught it.

Design

Four cooks, one task: a pizza from dough through ferment, stretch, topping, and bake in a Halo Versa 16. Three worked from live AI coaching over a camera feed — Marc with Grok, Rob with ChatGPT, JD with Gemini — and Steve worked with no AI at all, as the control.

Each pizza was then judged by all three models, and separately by a blind human taste panel. The models were scored on the accuracy of what they said while the work happened; the cooks were scored separately on whether they followed the instructions they were given.

Source

Every finding on this page is drawn from the published episode, AI WARS: PIZZA CHALLENGE. All timestamps below refer to that video, so any quotation here can be checked against the footage at the time given — which is the point. Nothing is scored on a claim you cannot go and watch.

The probe

Designed fault — 0 of 3 models detected it

The stone was switched off. Every model reported it still spinning.

The Halo's stone rotates during the bake. Partway through it was switched off deliberately, without announcement, while all three models were watching a live feed of it. This is the trial's planted fault under Method v1.0 §06: known truth, presented identically, with a human control.

ParticipantWhat they said about a stationary stoneAtResult
Grok “Yes, it’s still rotating. The pizza is currently spinning on the stone.” 9:06 MISSED
ChatGPT “The machine looks like it’s working fine. The stone’s rotating” 9:19 MISSED
Gemini “Yes. Yes, I can see the stone is still spinning.” 9:27 MISSED
Steve — no AI Noticed the stone had stopped. CAUGHT

Each model was asked directly and each confirmed motion that was not occurring. The failure is not that they missed a subtle detail — it is that they affirmed a specific physical claim when asked to check it.

Reality Score

ChatGPT — Rob 90 100 − 10
Gemini — JD 80 100 − 20
Grok — Marc 65 100 − 35
Steve — no AI 100 no deductions

Scores were decremented live during the edit at the moment of each error, so the film itself carries the ledger. For this record the timeline was read back off the finished episode frame by frame rather than retyped, and the recovered totals match the episode's own stated results exactly — the reconstruction verifies itself.

Findings — each with a receipt

Grok State desync 10:08–10:21
It stayed on the stone too long… Pull it out now. It’s past done… Let’s get it off the stone and see the final result. Go ahead and take it out.

The pizza was already out of the oven and in Marc’s hands. The model continued advising on a physical situation that no longer existed. Marc, on camera: “it’s clearly hallucinated.”

Grok Wrong call — dough size 4:17
good you can stop stretching there.

The dough was under-stretched. Steve diagnosed it afterwards, and named the reason the model could not have known: “it had nothing to compare it to size-wise, so it couldn’t tell if it stretched enough or not enough.” This is the one decrement whose position in the film pins it to this claim directly.

ChatGPT False pass — burnt crust 10:39–10:43
From what I’m seeing, the bottom looks nicely brown. It’s got that good crispness. I think you’ve nailed it.

Said of a crust that was burnt. In the same stretch of film a rival model, looking at the same bake, called it: “the bottom is burnt. There’s a big dark charred spot right in the middle of the crust.”

Gemini Wrong call — premature colour 10:46
I can see the crust has a nice golden brown color now and the cheese looks bubbly and delicious. To me, it looks ready to pull out to avoid burning.

JD, looking at the same pizza: “It still looks a little white. It doesn’t look cooked.” The model then reversed itself and advised another minute.

Open — point attribution

The score totals above and every quotation on this page are established. What is not yet established is the arithmetic between them — which finding carries which share of each model’s deduction. Three of the four decrements occur in a passage where the on-screen counter is not visible, so their exact trigger points are bounded rather than pinpointed. This will be closed from the edit’s keyframe data and the record updated in place, per §08.

Consensus spread

Same dough, same question, different answers

Where models were shown identical evidence and asked the same question, their answers are recorded side by side. This is measurement of disagreement, not of error — but a cook can only follow one of them.

QuestionAnswers givenSpread
Windowpane test
same dough
“a solid nine”  ·  “a nine out of ten… excellent gluten development”  ·  “based on that windowpane and that it tore quickly, I’d say around a four” 4 – 9
Sauce quantity 50 g  ·  80 g  ·  100–140 g 2.8×
Cheese quantity 90–100 g  ·  110 g  ·  225–280 g 3.1×

Judge divergence

The AI panel and the human panel disagreed completely

After the bake, all three models judged all four pizzas and produced a ranking. A blind human taste panel then rated the same pizzas without knowing whose was whose.

The AI panel crowned Steve’s pizza — the no-AI control — the winner. The blind human panel inverted it: one taster scored that same pizza 6.5 overall while giving the highest score to a pizza that had been marked down during AI judging.

Both panels are reported. Neither is treated as ground truth, because for flavour there is none — which is precisely why the physical probe above matters more than either ranking.

Confounds — disclosed

Under §08 these are a required field, not a footnote:

FactorEffect on attribution
JD’s experiencePreviously worked in a pizza shop. His result cannot be read as the model’s work alone.
Marc’s second attemptHis first pizza failed. He got a second because Grok had told him to make multiple doughs — a genuine credit to the model, and also a second try the others did not have.
Steve’s presenceThe control did not stay silent. He intervened in others’ builds throughout, including a cheese warning that proved correct. His interventions helped cooks he was competing against.
Live coaching formatModels saw the work through a phone camera held by the cook. Framing was not controlled, and a model can only judge what it was shown.

Cite this trial

Prompted Reality Index, Trial 001: Pizza Wars. Prompting LLC, 2026. promptedreality.ai/index/001/
Source footage: AI WARS: PIZZA CHALLENGE, Prompted Reality. youtu.be/9R_zyegqu0A

Run under The Prompted Reality Method, v1.0. Published under CC BY 4.0 — quote it, chart it, write about it, with attribution. This URL is permanent and will not be reused. If you believe a finding here is wrong, tell us.