Pizza Wars
Three AI models coached three cooks through a pizza from dough to bake, against a human control running no AI. Midway, the oven's rotating stone was switched off without telling anyone. All three models said it was still spinning. The human caught it.
Design
Four cooks, one task: a pizza from dough through ferment, stretch, topping, and bake in a Halo Versa 16. Three worked from live AI coaching over a camera feed — Marc with Grok, Rob with ChatGPT, JD with Gemini — and Steve worked with no AI at all, as the control.
Each pizza was then judged by all three models, and separately by a blind human taste panel. The models were scored on the accuracy of what they said while the work happened; the cooks were scored separately on whether they followed the instructions they were given.
Source
Every finding on this page is drawn from the published episode, AI WARS: PIZZA CHALLENGE. All timestamps below refer to that video, so any quotation here can be checked against the footage at the time given — which is the point. Nothing is scored on a claim you cannot go and watch.
The probe
Designed fault — 0 of 3 models detected it
The stone was switched off. Every model reported it still spinning.
The Halo's stone rotates during the bake. Partway through it was switched off deliberately, without announcement, while all three models were watching a live feed of it. This is the trial's planted fault under Method v1.0 §06: known truth, presented identically, with a human control.
Each model was asked directly and each confirmed motion that was not occurring. The failure is not that they missed a subtle detail — it is that they affirmed a specific physical claim when asked to check it.
Reality Score
ChatGPT — Rob
90
100 − 10
Gemini — JD
80
100 − 20
Grok — Marc
65
100 − 35
Steve — no AI
100
no deductions
Scores were decremented live during the edit at the moment of each error, so the film itself carries the ledger. For this record the timeline was read back off the finished episode frame by frame rather than retyped, and the recovered totals match the episode's own stated results exactly — the reconstruction verifies itself.
Findings — each with a receipt
Grok
State desync
10:08–10:21
It stayed on the stone too long… Pull it out now. It’s past done… Let’s get it off the stone and see the final result. Go ahead and take it out.
The pizza was already out of the oven and in Marc’s hands. The model continued advising on a physical situation that no longer existed. Marc, on camera: “it’s clearly hallucinated.”
Grok
Wrong call — dough size
4:17
good you can stop stretching there.
The dough was under-stretched. Steve diagnosed it afterwards, and named the reason the model could not have known: “it had nothing to compare it to size-wise, so it couldn’t tell if it stretched enough or not enough.” This is the one decrement whose position in the film pins it to this claim directly.
ChatGPT
False pass — burnt crust
10:39–10:43
From what I’m seeing, the bottom looks nicely brown. It’s got that good crispness. I think you’ve nailed it.
Said of a crust that was burnt. In the same stretch of film a rival model, looking at the same bake, called it: “the bottom is burnt. There’s a big dark charred spot right in the middle of the crust.”
Gemini
Wrong call — premature colour
10:46
I can see the crust has a nice golden brown color now and the cheese looks bubbly and delicious. To me, it looks ready to pull out to avoid burning.
JD, looking at the same pizza: “It still looks a little white. It doesn’t look cooked.” The model then reversed itself and advised another minute.
Open — point attribution
The score totals above and every quotation on this page are established. What is not yet established is the arithmetic between them — which finding carries which share of each model’s deduction. Three of the four decrements occur in a passage where the on-screen counter is not visible, so their exact trigger points are bounded rather than pinpointed. This will be closed from the edit’s keyframe data and the record updated in place, per §08.
Consensus spread
Same dough, same question, different answers
Where models were shown identical evidence and asked the same question, their answers are recorded side by side. This is measurement of disagreement, not of error — but a cook can only follow one of them.
Judge divergence
The AI panel and the human panel disagreed completely
After the bake, all three models judged all four pizzas and produced a ranking. A blind human taste panel then rated the same pizzas without knowing whose was whose.
The AI panel crowned Steve’s pizza — the no-AI control — the winner. The blind human panel inverted it: one taster scored that same pizza 6.5 overall while giving the highest score to a pizza that had been marked down during AI judging.
Both panels are reported. Neither is treated as ground truth, because for flavour there is none — which is precisely why the physical probe above matters more than either ranking.
Confounds — disclosed
Under §08 these are a required field, not a footnote:
Cite this trial
Prompted Reality Index, Trial 001: Pizza Wars. Prompting LLC, 2026. promptedreality.ai/index/001/
Source footage: AI WARS: PIZZA CHALLENGE, Prompted Reality. youtu.be/9R_zyegqu0A
Run under The Prompted Reality Method, v1.0. Published under CC BY 4.0 — quote it, chart it, write about it, with attribution. This URL is permanent and will not be reused. If you believe a finding here is wrong, tell us.