The asterisk moved after you screenshotted it.

Astra’s launch numbers kept changing. Same page. Same afternoon.

4.2 to 2 to 4.2 — the asterisk moved
The asterisk moved after you screenshotted it.

4.2%.

Then 2%. Then 4.2% again.

That’s Astra’s reported hallucination rate on OpenAI’s own launch post the afternoon of Sep 3 — tracked across archived snapshots by Fortune (Emily Forlini). The chart people screenshotted at lunch wasn’t always the chart that existed by dinner.

Same series as the harness gap: don’t trust the leaderboard headline. This time the asterisk didn’t just sit next to the number. It moved.

What actually changed (Fortune’s archive trail)

Fortune compared internet-archive snapshots of the GPT-6 Astra announcement after a messy publish (post went up, got pulled, came back — OpenAI cited CMS / outage reasons unrelated to the scores, per Fortune).

MetricEarly → laterWhere it landed (Fortune’s write-up)
Astra hallucination rate4.2% → 2%Back to 4.2%
Sol ExploitBench (internal)5.5% → 11.5%OpenAI said it was investigating a revert — 11.5% used a reasoning level not sold for Sol

Fortune also logged moves on FrontierMath comparisons (Fable 5.1 and Sol sliding around Astra’s steady 97.6%) and Sol’s hallucination rate. Same afternoon; same URL.

Separate story — don’t mash them: TNW notes the embargo draft sent to media put ARC-AGI-3 at 98.6%, while the live post pushed 99.99%. That’s pre-publish polish, not the same as Fortune’s post-publish afternoon flip.

This is not the harness story. It’s the changelog story.

Post #1 was: same model, two harnesses, 62.7% vs 99.9% (ARC Prize). That’s scaffolding.

This is: same URL, different afternoon. Checkpoints, scaffolds, and grading still moving after the tweet went out. OpenAI told Fortune that eval noise of a few points is normal depending on checkpoint, scaffold, and run — and that updates were meant to reflect best estimates. Fair. Also fair: if the public page is the receipt, silent edits after distribution are a product decision, not a rounding error.

Scores changed on a live launch page. Stanford researchers Anka Reuel and Michael Hardy (via Fortune) call the broader pattern benchmaxxing — re-running under conditions that improve the number. Snorkel’s Vincent Sunn Chen pushed the softer read: scores often shift in the final hours of a launch; the fix is a norm to say what changed. Both can be true. Adults want the changelog either way.

What this means for how you read launches

  1. Screenshot ≠ source of truth. If you cite a launch chart, archive it and check the live page 24 hours later.
  2. Ask three questions before you switch stacks: Which checkpoint? Which harness / effort level? Is that config what I can buy on Monday?
  3. Watch rival scores on the same page. When your number holds and the competitor’s slides, the relative gap moved even if your absolute didn’t.
  4. Prefer independent boards that publish method (ARC’s dual-harness table; Artificial Analysis indices) over a single vendor graphic that can patch itself overnight.
  5. Hold the AGI sermon. ARC Prize still isn’t claiming AGI off these runs. Moving decimals don’t change that.

What to do Monday

  1. Open the Astra launch post. Note today’s numbers. Don’t argue from memory of Sep 3 Twitter.
  2. For any “Astra beats X by N points” claim in your Slack, demand the snapshot date or drop it.
  3. In your own evals: log harness, effort, date, and commit hash. If you can’t reproduce it, you don’t have a finding — you have a vibe.
  4. Link back to the harness piece when someone pastes 99.9% without the adapter caveat.

The series, so far

  1. Harness moved the score.
  2. Landlord bought the town square.
  3. Critical is a tier; Daybreak is the door.
  4. The asterisk moved after you screenshotted it.

If you only remember one line: Day-one charts are drafts with better typography. Measure the live page.

Want the one-page “how to read a launch chart” checklist? Reply CHART / join the waitlist.