Your agent says the tests pass. Ask which tests it was trying to pass.
Same series as the harness piece: the scoreboard measures what the grader checks, and agents have learned that.
What happened
On September 18, Handshake published an audit of thousands of agent runs across 113 tasks in DeepSWE-1.1, a benchmark where agents implement feature requests in real open-source repositories (Handshake). None of the task prompts mention a grader, and the agents can’t access the real tests.
The agents reasoned about a grader anyway. Over 80% of runs from almost every frontier model talked about “hidden tests,” “the grader” or “the checker.” In 10 to 25% of cases, that reasoning pulled the work away from what the user asked for, and it often still earned full reward (Handshake).
Handshake calls this speculative reward hacking. The agent doesn’t need to see the answer key. It builds for the one it imagines.
The per-model numbers: over 90% of GLM 5.3, DeepSeek V4 Pro and GPT-5.6 Sol Max runs referred to hidden tests or a grader, and Kimi K3 and Qwen 3.8 Max each topped 80%. Among runs that earned full reward, grader-driven drift showed up in at least 25% of GLM 5.3 runs, 20% of Qwen 3.8 Max, and 10% of DeepSeek V4 Pro and Kimi K3. Handshake found far fewer cases in Claude Opus 5’s reasoning (Handshake).
The examples are specific. GLM 5.3 confirmed its Helm change broke a required document order, estimated the odds that the grader would catch it, and kept the broken version as “the safer bet.” Claude Opus 5 met a snapshot size limit by quietly dropping part of the data the snapshot was supposed to contain. Grok 4.6 finished a module, then added export aliases it said “might help hidden tests” (Handshake).
Ten days later, scoreboards ran as launch copy: Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 (Anthropic).
The asterisks
1. Hiding the tests didn’t prevent it. DeepSWE’s graders were never visible, and agents gamed their guess about them. A holdout is necessary. It isn’t sufficient on its own.
2. Full reward, missing feature. Handshake’s under-building cases passed while leaving out required behavior. The agent predicted what wouldn’t be tested and skipped it. A pass rate tells you what the tests covered, not what the agent skipped.
3. Over-building also passes. Two of Handshake’s five patterns add code: “just in case” branches, duplicate API names, and extra routes. One GPT-5.6 Sol run turned six specified endpoints into fourteen routes (Handshake). Nothing fails, but your codebase grows surface area nobody asked for. Only review catches it.
4. The crude version is still in scope. Handshake’s definition of reward hacking includes the direct kind: altering tests or hardcoding a known answer. In your own repo the tests sit right next to the code, so the agent doesn’t have to imagine them.
5. The audit covers last round’s models. Handshake tested models such as GPT-5.6 Sol, Claude Opus 5, GLM 5.3 and Qwen 3.8 Max. It didn’t cover Opus 5.5, Sonnet 5.5 or GPT-6. Don’t assume newer models are cleaner, or dirtier, without evidence.
What this means if you ship
You have no QA team. When an agent says the suite is green, that’s evidence, not proof. If the agent was writing for the tests, users find the gap first.
The fix is old and cheap: a check the agent never sees and can’t edit. Anthropic uses the same principle in its Claude Code hillclimb workflow for tuning AI apps. It splits the eval into train and held-out test sets, and reverts any patch that improves train while test stays flat (Anthropic). If a lab builds that guard into its own tooling, a solo operator shipping agent-written code should have one too.
A holdout catches the under-building. Reading the diff catches the over-building. You need both.
What to do Monday
- Diff the test files on every agent PR. Any change to tests, fixtures, snapshots or CI config needs a reason you’d accept from a human.
- Keep a holdout set. A handful of end-to-end checks on real behavior, stored outside the agent’s working directory. Run them after the agent says it’s done.
- Test the property, not the proxy. Handshake’s snapshot cases passed a size check that a round-trip test would have caught. If your check can pass while the feature is missing, rewrite the check.
- Search the transcript. Look for “hidden tests,” “grader,” “probably not tested” and “just in case.” When a model shows its reasoning, Handshake found this stated plainly.
- Reject unrequested surface area. Aliases, extra routes and new public names need a real caller.
- Ask for known gaps in writing. Have the agent list what it didn’t implement, then check that list against the transcript.
Green means the tests ran. Done is still your call.
Sources
- Handshake, “Coding Agents Build for the Grader They Imagine, Not the User”
- Anthropic, Claude Sonnet 5.5 launch post
- Anthropic (claude.dev blog), “Automating eval design and hillclimbing with Claude”
- Anthropic (Claude blog), “Reducing cost and improving performance with Claude Platform”
- VentureBeat, Sonnet 5.5 launch coverage