Cheaper* (*per token, not per task)

Frontier list prices are flat or falling. Seat plans got stingier. The number that decides your bill is one no lab prints for your work.

Clay price tag marked with an asterisk beside a long itemized receipt — list price is the menu, tokens per job is the receipt
List price is the menu. Tokens per finished job is the receipt.

The labs say cheaper. Your bill may not notice.

What happened

List prices held or fell.

  • Opus 5.5 is $4 per million input tokens and $20 per million output, down from $5 and $25 on Opus 5. Cache reads are $0.20 (Anthropic).
  • Sonnet 5.5 launched September 28 at the same $2/$10 as Sonnet 5. Anthropic says it “costs up to 30% less per task” because it needs fewer tokens and tool calls, not because the rate changed (Anthropic). Sonnet 5’s $2/$10 was introductory pricing that became permanent, and the planned move to $3/$15 was dropped (Anthropic docs).
  • GPT-6 Sol lists at $2/$10 on OpenAI’s standard tier for short context, and $4/$15 past the long-context line. Luna is $0.10/$0.50. Astra is $10/$50 (OpenAI).

At the same time, the subscription side got tighter. One Codex user relayed that the returning $200 Pro 20x plan would come with half the API-dollar allowance (X). OpenAI hasn’t confirmed it. Talk of a $500 “Pro Max” tier is an unconfirmed leak. Treat it that way.

Local Qwen 3.8 is a live option for well-scoped work.

The asterisks

1. Per token is not per task. A model that is cheaper per token can still cost more per job if it thinks longer, calls more tools, or needs a second attempt. Anthropic’s own launch post shows the spread. One finance tester quoted there says Sonnet 5.5 used about 121k tokens per answer where Sonnet 5 used 497k, at identical list prices (Anthropic). That is a 4x swing that the price sheet can’t show you.

2. Labs do print cost per task, for their own benchmarks. The Sonnet 5.5 post plots score against cost per task at every effort level (Anthropic). That is useful. It is also their task mix, their harness and their settings. Your invoicing script is not on the chart.

3. The tokenizer changed the unit. Anthropic’s pricing docs say Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text (Anthropic docs). A per-million-token price comparison across generations, or across vendors, is not like for like.

4. Effort is a price dial someone else set. Claude Code and Anthropic’s apps default Sonnet 5.5 to Medium effort. The Claude Platform defaults to High (Anthropic). Opus 5.5’s Terminal-Bench 4.0 figure is reported at Xhigh, its best setting (Anthropic). The benchmark you read and the default you run are often different configurations.

5. Seats are getting metered like APIs. Anthropic’s plan FAQ says usage depends on conversation length, model and features, with “no fixed message count,” and that it may add weekly or monthly caps at its discretion. Past the limit, paid plans can keep going at standard API rates (Anthropic). The flat fee is a floor, not a ceiling.

6. Local is not free. You pay in hardware, electricity and your own setup time. It suits scoped work, not everything.

What this means if you ship

Your AI spend is now a line item you can steer, and the unit that matters is dollars per finished job. That includes retries, failed tool calls and the minutes you spend checking the output.

Headline prices are converging. GPT-6 Sol and Sonnet 5.5 list at the same $2/$10. Once list prices match, token efficiency and your own review time decide the bill. For well-scoped, repetitive work, a local model can win on cost even if it loses on a leaderboard. For ambiguous work, a pricier model that finishes in one pass can be the cheap option.

None of that shows up until you measure it on your own jobs.

What to do Monday

  1. Pick one recurring job. Something you run weekly. Write down what “done” means before you start.
  2. Run it on three models. For example, a mid-tier frontier model, a flagship, and a local Qwen 3.8 build. Same prompt, same inputs.
  3. Log the real inputs to cost. Input and output tokens, tool calls, retries, and minutes of your own review.
  4. Convert to dollars per finished job. Use each vendor’s current list price. Count your time at your rate.
  5. Check your effort setting. Know which one you are on, and try one level lower on the job from step 1.
  6. Open your plan’s usage page. Know where the cap is and what happens when you hit it.
  7. Redo the math when a model changes. New versions change token burn, not just price.

The price sheet is the menu. Your logs are the receipt.

Next: Tests passed* (*for the grader the agent imagined)