Skip to main content

Tools / Free, runs on your machine

Claude Code Runway Audit

Find out where your Claude Code sessions spend tokens, and whether long sessions and compaction are why your usage runs out early in the week. One Python file. It reads token counts on your own machine and nothing else.

Run it

Save the file, then run it from any folder. It needs Python 3.8 or newer and nothing else.

python3 runway_audit.py
python3 runway_audit.py --days 14
python3 runway_audit.py --json > runway-baseline.json

Save the JSON output the first time you run it. Claude Code deletes old transcripts, so that file may be the only record you keep of how you worked before changing anything.

What the verdict means

  • LIKELY FIT: sessions of 40 or more turns carry at least half of your tokens, and you hit compaction at least once per 10 sessions.
  • POSSIBLE FIT: one of those two is true.
  • PROBABLY NOT A FIT: neither is true. Changing how you run sessions will probably not change much for you.
  • INSUFFICIENT DATA: fewer than 10 main sessions or 300 assistant turns in the window.

How it works

Claude Code writes one transcript file per session under ~/.claude/projects (or under CLAUDE_CONFIG_DIR if you set it). Each assistant turn records its token usage in four classes: uncached input, cache writes, cache reads and output. The audit adds those up per session, counts each message once even when it appears on several lines, and counts the compaction markers Claude Code writes when it summarizes a full context.

Subagent runs are counted separately from main sessions. Lines it cannot parse are skipped. If it finds no readable rows in the window it reports UNREADABLE with the versions it saw, never a row of zeros.

The fit rule is a screening rule, not a calibrated model. I chose its thresholds knowing my own history is heavy-tailed, so they may favor heavy users like me, and they have not been checked against anyone else. The rule was fixed before the result below was produced, and a test in the kit fails if it changes.

One real result: my own history

Window: 1 September to 28 September 2026. Denominators: 3,384 main sessions, 170 subagent runs, 42,783 assistant turns and 5,204,305,526 tokens, read from 5,120 transcript files written by Claude Code 2.1.281 to 2.1.285.

  • Cache reads were 92.49% of all tokens. Cache writes were 6.9%, output 0.59% and uncached input 0.02%.
  • Sessions of 40 or more turns carried 87.67% of main-session tokens.
  • The top 10% of main sessions (338) carried 92.04% of main-session tokens.
  • Main sessions hit compaction 878 times, 2.59 per 10 sessions. The median context size when compaction fired was 168,423 tokens.
  • Verdict: LIKELY FIT, for both reasons.

Limitation: this is one heavy user, and most of those sessions are automated one-shot runs (the median session has 2 turns), so the session count overstates interactive work. It describes where my tokens went. It is not a comparison and not evidence that any change saves tokens.

Free tools that measure the same data

Measurement is not scarce. Claude Code's own /context and /cost commands show the current session. ccusage reports usage across sessions from the same transcript files. claude-devtools shows compaction boundaries. This audit adds two things: the fit screen above, and a saved baseline you can compare against later.

If the verdict is LIKELY or POSSIBLE

The audit stays free. The paid Runway Kit adds the change steps: handoff and compaction templates, a way to trim the context every session loads, a status line that logs your plan usage, and a script that compares two periods of your own usage.

See what the Runway Kit adds

Questions

Does the audit send my data anywhere?
No. It is one Python file that uses only the standard library. It makes no network calls, needs no API key, and reads only token counts, model names, timestamps and compaction markers from the transcript files Claude Code already keeps on your machine. It never reads what you or Claude wrote.
Why does it say INSUFFICIENT DATA?
It needs at least 10 main sessions and 300 assistant turns in the window before it gives a verdict. Claude Code deletes old transcripts, so a short history can fall below that line. Widen the window with --days, or run it again after a few more weeks of work.
Is the fit verdict a prediction of how much I will save?
No. It is a screening rule that has not been tested against other people. It says whether your usage has the shape the change steps are aimed at: most tokens in long sessions, and regular compaction. It does not estimate a saving.
Does it work on Windows?
It has not been tested on Windows. It is written for macOS and Linux with Python 3.8 or newer.