How to Keep Cheaper AI From Blowing Up Your Budget
The Jevons paradox is running your AI bill. A thirty-minute cost-per-task audit tells you whether it's paying off.

So here's my confession. I pay $200 a month for Claude Max and another $200 for ChatGPT Pro, and my Grok plan is about to jump from $99 to $300. On top of that, my Hermes agent runs a local Qwen model and taps an OpenRouter account whenever it needs something bigger. Per token, the intelligence I'm buying keeps getting cheaper. My AI bill has never been higher.
I mean, I'd make fun of anybody else for this. And I've seen this movie. Back in the Cloud.com days, companies moved to the cloud to save money on servers, and then the first real cloud bill landed and the CFO wanted a meeting. Cheaper compute didn't shrink anybody's spend. It just made spinning things up so easy that nobody turned anything off.
There's a name for what's happening to your AI budget, and it's 161 years old.
Price the finished task before you celebrate cheaper tokens.
// The Takeaway: Falling token prices will grow your AI bill, because cheaper intelligence invites more work, bigger models, and longer agent runs. The number that tells you whether that's a good deal is cost per finished task: what you spend to produce one proposal, one resolved ticket, one research brief. Pull 30 days of spend, divide it by the work it produced, test a cheaper model on your priciest workflow, and set a hard spend limit. Give it thirty minutes and you'll know whether your AI budget is buying outcomes or just buying tokens.
Pull your last 30 days of AI spend → · Thirty minutes. Any paid API account (OpenAI, Claude Console, or OpenRouter). One spreadsheet.
In 1865, the economist William Stanley Jevons noticed something odd about Britain. Steam engines kept getting more fuel-efficient, and the country kept burning more coal. In The Coal Question, he called the assumption that efficiency means less consumption a confusion of ideas. "The very contrary is the truth." Cheaper steam power made steam worth using in places nobody had bothered with before.
Swap coal for tokens and you've got 2026. Bain & Company found that the average cost per token fell by half from December 2024 to December 2025, while tokens consumed grew 4.5 times over the same stretch. Fortune reported in June that token prices are down more than 90% since 2023, while spending on large language models has doubled since late 2025, according to the Silicon Data Token Expenditure Index. Same piece: Uber burned through its entire AI budget in the first four months of the year, then capped AI spending at $1,500 per employee per month.
// The real shift: Stop watching the price of a token and start watching the price of a finished piece of work. Cost per task is tokens per task times price per token, and tokens per task is the number that's quietly exploding.
Why the bill keeps climbing
Bain lays out three forces working against the headline price cuts. Nobody stays on last year's model; when a new frontier model ships, teams upgrade instead of pocketing the savings. Tokens per job keep climbing as agents take on multistep work, calling tools, fixing their own mistakes, and loading context. And once a team sees what agents can do, it finds ten more workflows to throw at them. Bain also found that the top 5% of users at a company often burn more tokens than the other 95% combined.
None of that is bad, by the way. That's what adoption looks like. The problem is spend that grows while nobody can say what it produced.
The good news is the biggest fix is usually model choice. Bain reports that AT&T, running 8 billion tokens a day, rebuilt its setup so large orchestrating agents hand work to smaller, domain-specific models, and says it cut costs 90% while tripling throughput. Bain also sees three-to-five-times cost differences between companies that match models to tasks and those that don't. You don't take the Ferrari to get groceries, right?
The thirty-minute cost-per-task audit
Pick your three busiest AI workflows and name the finished unit. Skip "writing help." Name the thing you ship: one first-draft proposal, one resolved Tier-1 ticket, one weekly competitor brief. If you can't count it, you can't price it.
Pull 30 days of spend, broken down by project, model, or key. In the OpenAI Platform, use the Usage page and break activity down by project (OpenAI Help Center). In the Claude Console, open the Cost page, pick the workspace, filter by model, and export to CSV (Claude Help Center). On OpenRouter, go to openrouter.ai/activity, group by Model or API Key, and choose Export to CSV from the options menu (OpenRouter docs). If every workflow shares one key, that's your first fix: give each its own project or key.
Count the finished units for the same 30 days. Pull them from your CRM, your help desk, or your sent folder. A rough count beats no count.
Divide. Workflow spend ÷ finished units = cost per task. Put it next to what the task costs a human (loaded hourly rate × the hours it used to take). Now you know whether that workflow's growth is earning its keep.
Run a downgrade test on the priciest workflow. Take five real inputs from last week, run them through the next-cheaper model, and grade the results against what you actually shipped. If four out of five would have gone out the door, switch that workflow and save the frontier model for work that needs it.
Set an alert and a hard cap. In the OpenAI Platform, go to Organization limits, select Edit spend limit under Spend, and turn on Enforce a hard limit; each project can get its own limit too (OpenAI docs). Alerts only notify; a hard limit stops requests, and enforcement isn't instant, so set the alert well below the cap (OpenAI Help Center). In the Claude Console, open a workspace and use its spend limits to cap monthly spend and set alert thresholds (Claude docs). On OpenRouter, give each agent its own API key with a credit limit that resets daily, weekly, or monthly (OpenRouter docs).
Put fifteen minutes on the calendar every month to redo step 4. Track cost per task, not the total bill. A rising bill with falling cost per task is a business that's scaling. A rising bill with rising cost per task is a leak.
Bonus: your team is on seats, not API keys
Seats feel flat-rate right up until they aren't. On Claude Enterprise, admins can set per-user spend limits, with defaults by seat tier, group, or the whole organization, all configured in claude.ai (Claude API reference). If your developers sign in to Claude Code through Console billing, that Claude Code workspace is the only Console workspace that supports per-user monthly spend limits (Claude docs). Uber's flat $1,500 cap is the blunt version of this move. The better version sets each team's cap from the audit above, run per team instead of per workflow.
Jevons isn't the villain
Here's the part I like. The same paradox running up your bill is the best counterargument to the "AI eats every job" crowd. Apollo chief economist Torsten Slok points out that call center employment in the Philippines has nearly doubled over the last decade, even as AI got good at customer service work, because a cheaper interaction means more customers served (Fortune). Cheaper intelligence means more work worth doing.
So I'm not going to tell you to spend less on AI. I'll probably spend more next year, and so will you. Just know what you're buying. The companies that come out ahead won't be the ones with the smallest token bill. They'll be the ones who can tell you, to the penny, what a finished proposal costs.

Your AI Sherpa,
Mark R. Hinkle
Founding Publisher, The AIE Network
Follow me on LinkedIn
