I am token-addicted. Not in a cute, ironic way — in the “ideas keep turning into real things whilst I’m selling my time for money” way.
Billable work is running. A client call. A review. Something sensible. And next to it, a second stream of prompts is chewing through Claude, Copilot, and Cursor like snacks at a desk that never empties. Side projects that used to live in a notes app for years now ship in an afternoon. That part is genuinely exciting. The bill is less exciting.
Three subscriptions, one habit
I don’t have one AI subscription. I have Claude, Copilot, and Cursor. Each one resets on its own schedule. Each one feels wasted if I leave capacity on the table. So when a week is almost over, I get this little itch: don’t let the included tokens expire unused. That’s not productivity. That’s sunk-cost theatre with a usage meter.
You know the feeling if you’ve ever stared at a gym membership you barely used and then gone three times in the last four days of the month. Same energy. Worse UI.
A menu bar app, because denial is expensive
I use ClaudeBar — a free, open-source macOS menu bar app that tracks usage quotas across Claude, Copilot, and friends. It doesn’t fix the habit. It just makes the burn rate visible enough that I can’t pretend it’s “only a few prompts”.

Visibility is useful. It also turns every glance into a tiny hit of dopamine or mild panic, depending on the day.
Smaller models, and a local escape hatch
I’ve started leaning on smaller models more often — Sonnet instead of the biggest thing on the menu. The pitch I tell myself: lower cost per prompt, supposedly more throughput for the same quota. Whether that’s true or just coping is still an open question. It feels like I get further before ClaudeBar goes orange.
At home there’s another trick. For analysing private documents I run a local LLM with Ollama. Same desk, different meter: electricity instead of tokens. Nothing leaves the machine, the quota doesn’t flinch, and the only thing that suffers is the power bill. For personal paperwork that trade-off is an easy sell.
The awkward bit remains. The more I use these tools for real work and for the idea pile, the harder it is to tell which spend is leverage and which is just filling the quota. Shipping faster is real. So is the urge to burn what’s “already paid for”.
What happens when the subsidy ends?
Here’s the bit that keeps me up longer than the usage meter: I fear that autumn 2026 is when the market consolidates, and the heavily subsidised tokens I’ve been mainlining get repriced into something closer to their real cost.
The pattern is already visible. Flat-rate plans have been acting like discounted token bundles for power users — people burning API-equivalent value far above what they pay.³ Labs are still losing money at a scale that only makes sense in a land-grab phase.¹ ² And vendors are already shifting from soft unlimited vibes toward usage-based billing as inference bills climb.⁴
If that squeeze lands hard, my “included tokens” habit stops being a quirk and starts being a budget problem. Do I push more work onto local models? Ration harder? Or just pay the new sticker price and pretend it’s fine?
I don’t have a clean answer yet. I know the stack is expensive. I know the side ideas becoming reality is the best part of this moment in software. And I know “use it or lose it” is a terrible way to decide what deserves attention — especially if “lose it” soon means “couldn’t afford it anyway”.
Are you preparing for the subsidy cliff — or still racing the reset like nothing will change?