Usage* (*the agent re-reads your whole chat, every turn)

Anthropic’s own docs say Claude Code resends your full conversation on every request. Here’s what that costs and how to see it in your own sessions.

Six paper stacks rising left to right, each topped with a clay sheet, an asterisk over the tallest
You pay for the question. You also pay for everything said before it.

Your usage meter counts tokens. It does not tell you how many of them were the agent reading the same conversation again.

What happened

Anthropic’s Claude Code cost docs explain how a session is billed: “Claude Code sends your full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results.” With prompt caching, Claude Code “re-reads that history at the cached token rate, so a one-line question in a session that has been open all day still draws usage for the whole conversation” (Anthropic docs).

Caching makes each re-read cheap, not free. Anthropic’s pricing page lists a cache hit at 0.1x the base input price for most models, 0.05x on Opus 5.5, and 0.025x on Fable 5.1 and Mythos 5.1 (Anthropic docs). A small fraction of a growing number is still a growing number.

The docs list other ways a session resends its whole context. A cache miss after a break “reprocesses your full context.” The cache lifetime is an hour on a subscription and five minutes by default on an API key. Scheduled tasks and idle check-ins can also send your full context while you are away (Anthropic docs).

Anthropic also names the usual cause of a surprise bill. On API and cloud plans, it says high spend “usually traces back to long sessions that were never cleared or to Opus left as the default model” (Anthropic docs).

The asterisks

1. Cheap per token is not cheap per session. A cache read costs a fraction of base input, but a long session pays that fraction again on every turn and every tool call. Discounts on repeated tokens reward repetition. They don’t prevent it.

2. Resetting has costs too. The docs note that /compact “reads the conversation it summarizes,” so compacting a large context is itself a large request. /clear costs nothing, but it drops the history you may still need.

3. Subagents move tokens, they don’t erase them. A subagent does its work in its own context and returns a summary, which keeps the main conversation smaller. Anthropic notes it “also sends its own requests, which count toward the same usage limits” (Anthropic docs).

4. The dashboard number is an estimate. Claude Code computes the dollar figure in /usage locally at list price, and the docs point to the Claude Console usage page for authoritative billing. On Pro and Max plans, the session cost figure isn’t what you are billed (Anthropic docs).

What this means if you ship

For a solo operator, the model on the price sheet is only part of the bill. How long a session runs before you reset it may matter as much. A session that carries three finished tasks into a fourth pays to re-read all three, every turn.

That is a habit problem more than a vendor problem, which is useful news. You can change habits by Monday. The catch is that you can’t manage what you haven’t measured.

What to do Monday

  1. Read your own cache line. In recent versions, /usage shows a Prompt cache (main) line with the share of input tokens served from cache, plus misses (Anthropic docs). A high share in a long session tells you most of your input is history being re-read. It covers the main conversation only, not subagents.
  2. Check the breakdown. On Pro, Max, Team and Enterprise plans, /usage attributes usage to subagents, skills, plugins and MCP servers, and flags behaviors such as long context or cache misses at 10% or more of recent usage (Anthropic docs).
  3. Run /clear between unrelated tasks. Anthropic’s advice: “Stale context wastes tokens on every subsequent message.” Use /rename first so you can /resume later.
  4. Compact earlier, with instructions. /compact Focus on code samples and API usage tells Claude what to keep. Compacting a small context is a smaller request than compacting a huge one.
  5. Send noisy work to subagents. Tests, documentation lookups and log processing can stay in the subagent’s context while only a summary comes back (Anthropic docs).
  6. Keep a short handoff notes file. CLAUDE.md files load at the start of every session, so a fresh session can start with what matters. Anthropic suggests keeping each under 200 lines, because longer files consume more context (Anthropic docs).

The meter shows what you spent. The transcript shows why.

Next: Safe* (*safe from whom?)