// AI Advantage
Your Top 5% Is Your AI Budget
A handful of people are spending most of your AI money. Find out whether they're your best operators or your biggest leak.

Your AI budget probably belongs to five people, not five hundred. Bain & Company found that the top 5% of users at a company often burn more tokens than the other 95% combined. I'll admit I'm one of them. Yesterday I confessed to paying for three premium AI plans, and I still run into limits most weeks.
That lopsided split is the most useful fact about your AI spend, and averages hide it. So does a per-seat line item. When a handful of people drive most of the bill, there are two explanations, and you want to know which one you've got.
Either those people are your best operators, the ones who figured out how to hand real work to AI, and you should get them in a room and copy what they do. Or they're running a leaky workflow: an agent looping on the same task, a 200-page PDF re-uploaded into every chat, the biggest model doing grammar checks. Your plan's admin console or your API usage page will tell you who they are. Fifteen minutes over coffee will tell you which kind.
Usually it's both. Your power users do great work and waste a surprising share of their budget doing it. Five fixes stretch a plan without slowing anybody down:
Start a new chat for each new task. Anthropic lists conversation length as one of the things that eats your limits, since every reply works from the whole thread. That three-hour chat is expensive for a reason.
Put the documents you reuse into a project. In Claude, project content is cached and counts less against your limits when you reuse it. Upload the style guide once instead of pasting it into every chat.
Ask everything at once. Batch related questions into one message, and edit your prompt instead of sending a correction. Every round trip costs you.
Match the model and the tools to the job. Model choice, Research, and web search all affect how fast you hit limits. Save the biggest model and deep research for work that needs them.
Queue the API work that can wait. If a job doesn't need an answer this hour, Anthropic's Message Batches API runs it at half the price.
None of this is about spending less. Your top 5% should probably get more budget, not less. They should just spend it on work instead of on rereading the same thread.
Run the usage report this week. Find your five people, buy them lunch, and ask two questions: what are you using this for, and what would break if we capped it? Then hit reply and tell me what you found.
Your AI Sherpa,

Mark R. Hinkle
Founding Publisher, The AIE Network
Follow me on LinkedIn
If you want to get in contact or give me feedback, just reply to this email. I read every single one of them.
