4.2%.
Then 2%. Then 4.2% again.
That’s Astra’s reported hallucination rate on OpenAI’s own launch post the afternoon of Sep 3 — tracked across archived snapshots by Fortune (Emily Forlini). The chart people screenshotted at lunch wasn’t always the chart that existed by dinner.
Same series as the harness gap: don’t trust the leaderboard headline. This time the asterisk didn’t just sit next to the number. It moved.
What actually changed (Fortune’s archive trail)
Fortune compared internet-archive snapshots of the GPT-6 Astra announcement after a messy publish (post went up, got pulled, came back — OpenAI cited CMS / outage reasons unrelated to the scores, per Fortune).
| Metric | Early → later | Where it landed (Fortune’s write-up) |
|---|---|---|
| Astra hallucination rate | 4.2% → 2% | Back to 4.2% |
| Sol ExploitBench (internal) | 5.5% → 11.5% | OpenAI said it was investigating a revert — 11.5% used a reasoning level not sold for Sol |
Fortune also logged moves on FrontierMath comparisons (Fable 5.1 and Sol sliding around Astra’s steady 97.6%) and Sol’s hallucination rate. Same afternoon; same URL.
Separate story — don’t mash them: TNW notes the embargo draft sent to media put ARC-AGI-3 at 98.6%, while the live post pushed 99.99%. That’s pre-publish polish, not the same as Fortune’s post-publish afternoon flip.
This is not the harness story. It’s the changelog story.
Post #1 was: same model, two harnesses, 62.7% vs 99.9% (ARC Prize). That’s scaffolding.
This is: same URL, different afternoon. Checkpoints, scaffolds, and grading still moving after the tweet went out. OpenAI told Fortune that eval noise of a few points is normal depending on checkpoint, scaffold, and run — and that updates were meant to reflect best estimates. Fair. Also fair: if the public page is the receipt, silent edits after distribution are a product decision, not a rounding error.
Scores changed on a live launch page. Stanford researchers Anka Reuel and Michael Hardy (via Fortune) call the broader pattern benchmaxxing — re-running under conditions that improve the number. Snorkel’s Vincent Sunn Chen pushed the softer read: scores often shift in the final hours of a launch; the fix is a norm to say what changed. Both can be true. Adults want the changelog either way.
What this means for how you read launches
- Screenshot ≠ source of truth. If you cite a launch chart, archive it and check the live page 24 hours later.
- Ask three questions before you switch stacks: Which checkpoint? Which harness / effort level? Is that config what I can buy on Monday?
- Watch rival scores on the same page. When your number holds and the competitor’s slides, the relative gap moved even if your absolute didn’t.
- Prefer independent boards that publish method (ARC’s dual-harness table; Artificial Analysis indices) over a single vendor graphic that can patch itself overnight.
- Hold the AGI sermon. ARC Prize still isn’t claiming AGI off these runs. Moving decimals don’t change that.
What to do Monday
- Open the Astra launch post. Note today’s numbers. Don’t argue from memory of Sep 3 Twitter.
- For any “Astra beats X by N points” claim in your Slack, demand the snapshot date or drop it.
- In your own evals: log harness, effort, date, and commit hash. If you can’t reproduce it, you don’t have a finding — you have a vibe.
- Link back to the harness piece when someone pastes 99.9% without the adapter caveat.
The series, so far
- Harness moved the score.
- Landlord bought the town square.
- Critical is a tier; Daybreak is the door.
- The asterisk moved after you screenshotted it.
If you only remember one line: Day-one charts are drafts with better typography. Measure the live page.
Sources
- Fortune — Astra eval metrics changed post-launch
- TNW — harness gap + metric movement wrap
- ARC Prize — Astra on ARC-AGI-3
- Series: #1 · #2 · #3
Want the one-page “how to read a launch chart” checklist? Reply CHART / join the waitlist.