|
A field guide You probably don’t need the bigger plan.Ten things to check before you upgrade, in the order I’d check them. Most people running out of tokens are not doing too much work. They are carrying too much furniture into every single message. I’ve built two seven-figure software companies, and I shipped one of my mobile apps on a single $20 subscription. So when someone tells me they’re hitting limits by lunchtime and the fix is a plan four times the price, my first question isn’t how much do you code. It’s what are you dragging along behind you. Almost always the answer is: a lot. And almost always they have never looked. The mechanic everything else follows from These tools are stateless. Every time you hit enter, the entire conversation is sent again from scratch — your prompts, every response, every file it read, every tool result, plus your instruction files, your skill definitions, and the description of every MCP tool you have installed. It is not reading a log or remembering where it left off. It is re-processing the whole pile, on every turn. So a 2,000-token file you loaded forty messages ago isn’t a one-time cost of 2,000 tokens. It is 2,000 tokens times every message you’ve sent since. Caching softens this and it does not remove it. On Anthropic’s published API pricing a cached token still bills at roughly 10% of the normal input rate, and writing to the cache costs about 25% more than a plain token the first time. Subscription plans don’t publish their internal weighting, so I won’t pretend to quote you a number there. But the shape holds: cheap is not free, and 10% of an enormous number is still a large number. 10% What a cached input token still costs on the API. Discounted, not free. 5× Input-price gap between Haiku 4.5 and Opus 5. Same dollar, five times the volume. Every turn How often your skills, instructions and tool list get resent. Not once. Every turn. ~30 min The one-time audit in Part two. It’s the highest-return half hour here. Everything below applies whether you’re on Claude Code or Codex, since the underlying economics are the same. Where a command or file path is specific to Claude Code I’ve said so, and I’ve checked each one against a current install rather than repeating what circulates on social. Two items that show up on nearly every list like this one are wrong; I’ve flagged both where they come up. What’s in here
01
The audit. Three things to measure and cut before you change a single habit. This is where the money is.
02
The habits. Five working patterns that stop refilling the window you just emptied.
03
The two commands. Both underused, one of them almost unknown.
04
What doesn’t work. Including the trick half the internet recommends, and why I don’t.
05
If you’ve done all ten. When a second subscription genuinely beats a bigger one. Part one · The audit Three things you are paying for on every message, whether you use them or not.This is a one-time cleanup with a permanent payoff, which makes it the only part of this document I’d call urgent. Do it once and every conversation you have afterwards starts lighter.
01
Audit your skills, and tighten every trigger.This is the single biggest hidden cost I see, and it’s gotten worse as skill marketplaces filled up. You installed thirty of them over a few months. You actively use four. The other twenty-six are still being described to the model on every turn so it can decide whether they’re relevant. A well-built skill costs you around a hundred tokens to have sitting there — a name and a tight description, with the real content loaded only when it fires. A badly-built one can pull in thousands before you’ve seen a single word of response. The worst offenders are consistently the “my voice” and “content coach” genre, because their trigger conditions are written so broadly that the model concludes it needs the full skill loaded to shape almost any output. You ask a one-line question about a config file and you’ve paid for a 4,000-token essay on your brand tone. So go through them. For each one, read the trigger description and ask whether it fires exactly when you’d expect and never otherwise. Tighten anything vague. Delete anything you haven’t used in a month — you can always reinstall it. Progressive loading is a genuinely good design, but if a skill progressively loads on every turn, you are paying full price for a feature you think is saving you money. Where to look
The test Would this fire on “fix the typo on line 12”? If yes, the trigger is too broad. Effort Fifteen minutes, once. The largest single win on this list.
02
Put CLAUDE.md on a diet, then split it up.
Your instruction file — The specific way these explode is worth naming, because almost nobody notices it happening. Somebody sets up a rule along the lines of “any time I correct you, write it down in CLAUDE.md so you don’t do it again.” That sounds like excellent hygiene. Six weeks later the file is nine hundred lines of accumulated micro-corrections, most of them contradictory, most of them about a file you deleted, and all of them being resent on every single turn. The fix is two moves. First, cut it hard: a short, sharp root file with the things that are true everywhere — how to run the tests, the conventions that actually matter, what not to touch. Second, push the specifics down. Both Claude Code and Codex will pick up instruction files in subdirectories and load them only when work moves into that directory, which is how you get detail without paying for it constantly. Frontend conventions live in the frontend folder. They do not need to be in context while you’re editing a database migration. The number you’ve probably seen, with the caveat attached A developer write-up that circulated widely reported cutting always-loaded context from 42,200 tokens to 1,900 — a 94% reduction — and extrapolated that to roughly 1.2 million tokens a day at thirty conversations. I’ve seen no independent verification of either figure and the daily number assumes a very particular usage pattern, so treat it as an illustration of the shape rather than a benchmark. The shape is right regardless: a bloated instruction file is a fixed tax multiplied by every message you will ever send. Target Root file under ~100 lines. If it’s longer, something belongs in a subdirectory. Watch for Auto-append rules. They’re the mechanism that inflates the file invisibly. Effort Ten minutes, plus a quarterly re-read.
03
Run
|
| The lever | Honest size | Why it lands where it does |
|---|---|---|
| 01 · Skill audit | Large, permanent | Fixed cost on every turn, and the one most people have never looked at |
| 02 · CLAUDE.md diet | Large, permanent | Same mechanic, and these files inflate silently over months |
03 · /context |
Zero on its own | Saves nothing directly. Tells you which of the other nine is your problem. |
| 04 · Model matching | Large, ongoing | Up to a 5× input spread. Requires you to actually switch, which is the hard part. |
| 05 · MCP pruning | Medium to large | Scales with how much of a collector you’ve been |
| 06 · Point at the file | Medium, compounding | Avoided reads never enter the window, so they’re never resent |
| 07 · Deny rules | Small, occasionally huge | Insurance. Most days nothing; the day it catches a bundle it pays for itself |
| 08 · Session hygiene | Medium, ongoing | Hardest to make habitual, which is why it’s usually the one still undone |
| 09 · Slash commands | Small directly | Pays through determinism. Retries cost far more than verbosity does. |
10 · /btw |
Small to medium | Keeps side questions out of the permanent history, and out of the model’s way |
| Output compression | Marginal, possibly negative | Targets the small side of the ledger and buys you extra rounds of clarification |
Part five · If you’ve done all ten
Then the honest answer might be two subscriptions, not a bigger one.
This is the mildly controversial part, and it only applies once the ten above are genuinely done rather than skimmed.
If you’ve run the audit, cut the files, pruned the servers, and you are still hitting limits doing real work, then you have a capacity problem rather than a hygiene problem, and it’s worth solving. But before you quadruple your spend on a single provider, do the arithmetic on a second $20 subscription instead — either a second one with your provider of choice, or one with a different provider entirely.
Two independent limits at $40 is often more usable capacity than one larger limit at a higher price, and it buys you something a bigger plan can’t: a second opinion. Different models fail differently. When one is stuck in a loop on a bug, handing the same problem to a different family frequently unsticks it in one turn, which is a token saving in its own right.
The thing that makes or breaks this is switching friction. If moving between models means reconfiguring your environment, you will not do it, and you’ll have bought a subscription you use twice. Pick a harness that makes switching trivial — conductor.build is what I use, and it also lets you run several agents in parallel, which changes the calculus again. Whatever you pick, test the switching flow before you commit to the second subscription.
The order I’d actually do this in
Run /context in a fresh session and write the number down. Audit skills, then
CLAUDE.md, then MCP servers. Run /context again and compare. Work for a week
pointing directly at files and switching models by task. If you are still constrained after that, and
only then, buy the second subscription — at which point you’ll know precisely what you’re buying and
why, instead of guessing.
Running out of tokens feels like a capacity problem, which is why the pitch for a bigger plan lands so easily. Most of the time it’s a carrying problem. You are moving the same furniture into every room and being surprised that it’s heavy.
The whole audit takes about half an hour and costs nothing. A plan upgrade takes thirty seconds and costs you every month from here on. It seems worth doing them in that order.
One more thing
If you run the audit and the number still doesn’t make sense.
Sometimes the floor is high for a reason that isn’t on this list, and it’s hard to spot from inside your own setup.
Send me your /context breakdown from a fresh session — just the numbers, no code, nothing
proprietary — along with roughly what you’re building and which plan you’re on. I’ll tell you which of the
ten above is eating your budget and whether upgrading would actually fix it. Often it wouldn’t, which is a
cheaper answer than finding out in a month.
Get in touch
Email hi@davecto.com with the subject line “Token Diet” and paste the breakdown.
More guides like this one, for people trying to use AI without embarrassing themselves.
Weekly, plain-language breakdowns on Instagram.
@davecto
A note on sourcing: per-token prices and the caching multipliers are Anthropic’s published API rates as of
writing, and they describe API billing rather than subscription-plan limits, which aren’t published in
comparable terms — the ratios are the useful part, not the dollar figures. Commands, file paths and settings
syntax were checked against a current Claude Code install (2.1.x) rather than taken from secondary write-ups;
that is also how I established that .claudeignore is not a real feature, and it’s worth
re-checking after major releases since this surface moves. The 42,200-to-1,900-token figure is one
developer’s self-reported result, repeated here with that caveat and not independently verified. The
relative sizing in the Part four table is my judgement from my own usage, not measurement across a
population — your /context output is better evidence about your setup than my table is. I have
no commercial relationship with Anthropic or OpenAI. I do use Conductor and recommend it on that basis;
it’s named because switching friction is the thing that decides whether a second subscription pays, not
because anyone paid for the mention.