Tackling your board's next big question

The next great cost problem

Aug 14 | 6 min read | By Tim Cooper

TLDR;

AI costs are spiraling, as usage based billing catches CFOs off guard. Suddenly the shiny new toy looks like the angry Little Shop of Horrors plant demanding “Feed me [tokens] now!” Visibility from many providers sucks, making it hard to contain the problem and control token efficiency. And CFOs are having to adapt fast:

  • Model menu: Understand the cost differential and suitability of each model for different tasks. Educate the business on model routing and optimization.

  • Measure yourself: Set up internal token cost tracking by employee, model and other variables.

  • Set limits: Implement hard guardrails, to avoid embarrassing mistakes.  

Someone in the board meeting asks why DSO ticked up. Or which three accounts are quietly wrecking the aged debt report…  

You give the scripted nod that means you’ll "circle back ASAP" which everyone in the room correctly reads as "I don't know."

Here's the really annoying part: the answer was sitting in your data the whole time. It just needed a translator.

In an alternate reality, Stuut's AI agent has been running your collections for a while now. The outreach, the chasing, the cash actually landing. It always had the receipts. Now you can just ask it things, in plain English, and it answers in seconds (charts, tables, invoice numbers included) direct from your live ledger. 

You’ve got the answer before someone asks the question.

A year ago, we were talking about headcount savings from AI and didn’t much consider the cost.

A year is a long time in AI.

There is very little evidence of headcount savings demonstrably attributable to AI…yet. Although many high growth companies have claimed growing faster or doing more “with the same.” 

Meanwhile, the conversation has shifted from replacing people to running up the bill. AI spend has exploded. Anthropic’s annual revenue run rate grew to $47 billion in May (from zero three years ago). And we can assume most of that is corporate spend on AI.

The boom in AI spend is showing up in other places as well. AI token spend for customers of the financial ops platform Ramp have increased 20.7 times since June 2025. At this pace, any assumptions about future AI token cost should be tossed. 

So, token spend has quickly become the hottest topic, especially among tech businesses.

But first, what is a token?

It’s a unit of intelligence in AI that LLMs use to charge for usage. One token typically means three to four characters of input or output from an AI model. As a rough guide, 100 tokens works out to about 75 words, give or take, depending on the language and which model is counting. 

Providers such as Anthropic, OpenAI, and others typically bill input and output token differently, with output usually costing more. The variables are the volume of tokens consumed, and the cost of an individual token. Different AI models, not just model providers, are priced differently on a per token basis. For example, Claude's Fable 5 is twice the price per token compared to its Opus 5 model.

So, what model you use, the price of a token, and how many tokens you consume is important.

Some models price on a fixed subscription but with usage restrictions. That could be a cap on the tokens you can consume before you get pushed into another package. For example, after reaching their use limits, subscribers to Claude plans (Pro, Max 5x, and Max 20x) can switch to consumption-based pricing at standard rates, so they can continue working.

Users who are coding or producing other serious output from AI models will blow through these limits quickly and move to usage-based pricing.

Some models have even done away with these caps and moved to a full usage base. These pricing changes are tricky for CFOs because:

  1. Tokens are an unfamiliar currency - the value you get from 1,000 tokens varies wildly.

  2. The cost of a token varies based on model and package over time

  3. We have no or limited transparency over how tokens are used. Even if we did have a good view, how do we measure the value of what they were used for?

We are learning this all in real time, meanwhile we’ve seen a few horror story headlines of folks who didn’t do (or have) the math right:

  • An unnamed tech company got stuck with a shock $500m Anthropic bill in one month after it failed to set token use limits for employees.

  • Uber used up its 2026 AI budget by April. Uber COO Andrew Macdonald called it a “head exploding moment,” sparking conversations about headcount trade-offs.

  • Amazon blew $1.8 million on a Claude project to match author’s names with online listings. The spend was 860% over budget and no one caught the f*$k up for five months. 

These price shocks are largely driven by AI model providers switching from per seat to use-based charges with no meaningful guardrails.

It’s not just big companies – this issue is everywhere. One consultant told us: “CFOs are saying: ‘we burnt our entire year’s token budget in two days – I didn't know all these people were using it, what for, or what’s the ROI.’”

What’s driving higher spend?

Ramp said the biggest driver of unexpected cost increases is teams upgrading from light to frontier versions for quality, often without realizing that could hike prices 10 or even 100 times. Premium model cost share rose from 5.7% in June 2025 to 55.9% in April 2026.

Also models and prompts used have grown; and organizations are using more agentic AI, which can eat tokens at alarming rates.

Samantha Greenberg, CFO of AI market intelligence platform AlphaSense, said the situation has changed radically within six months. “Realizing AI use costs can rise quickly is frightening for CFOs. Some are reacting by constraining use. A nuanced approach avoids slamming the brakes on AI for productivity and innovation but instead focuses on optimizing token use.”

She said firms are laser-focused on token efficiency strategies, such as:

  • Routing the lightest model necessary for each job or function, not just defaulting to the most convenient or capable

  • Training users on efficiency

  • Customizing harnesses (the supporting systems around LLM models)

  • Pushing model diversity and avoiding lock-in contracts.

Greenberg said: “We’ve spent enormous time and effort on training token optimization, and routing which models to use when. Then we review things like:

  • Training our own models

  • Using simpler, more token efficient workflows.”

AlphaSense has also built its own harness intelligence around the LLM. Greenberg said building that has been top of mind since an MIT / Stanford report from March 2026 showed a finely tuned harness can improve performance sixfold.

 

Sharpening visibility

Another reason many CFOs are flipping out when they see their Anthropic bills is poor visibility.

Ramp said: “Unlike seat-based software, token spend scales with use, and can grow quickly across teams. It often sits behind provider dashboards and invoices that are difficult for finance teams to interpret. That makes it harder to track spend against plan, identify inefficient use, and understand where AI investment is going.”

Ramp aims to fix this by surfacing savings opportunities, and giving finance teams the context they need to optimize spend before it grows unchecked.

Amy Wang, CFO at procurement platform Procurify, said: “We’ve had our own surprises. For example, ‘you’ve consumed over 80% of token allotment in the third quarter. It’s time to buy more credits.’ That was a concern as those users were about to ramp up consumption further.”

Asking providers for reports about token use and options before the bill arrives had limited success, she said.

“But this is a big spend for us and it can get out of hand within an hour or a day. Even last week, we saw a 10x spike when everyone was trying to query a specific problem in the business. So we’ve built an internal tool to monitor our use,” said Wang.

Procurify is also:

  • Reviewing its AI stack to understand how it’s charged and measured

  • Monitoring consumption closely – Wang oversees a rev-ops team, which helps her track go-to-market and R&D tool use early

  • Setting limits on accounts and dialing back any tool not delivering expected ROI

  • Developing internal databases to reduce external costs.

“Many companies are building their own context protocols [which connect AI apps to their software] to avoid consuming credits through external tools. Through all these efforts, we’ve reduced token use. It’s hard to measure but we aim to avoid our bill being two or three times over budget,” said Wang.

Greenberg said poor visibility is an opportunity as well as a pain. “Finance can add value by embedding data scientists with FP&A to navigate to real-time insights and forecasts. That drives the visibility the organization needs to increase token efficiency. I can see platform use at any day or time, for each customer, and each revenue vertical.”

She also recommended this combined team constantly communicate with go-to-market, customer success, and product teams.

“[For tech firms], give your customers visibility with dashboarding tools so they understand their use and cost impacts. Every customer uses your platform differently. The more we can learn about use patterns around factors such as sector, geography and role, the better the forecast will be,” said Greenberg.

More efficiency tips

Some companies are also:

  • Setting up centers of excellence or super user groups to streamline AI use across the enterprise

  • Negotiating price per business outcome, such as workflow forecast or report, rather than per use

  • Hiring fin-ops roles, and finance talent with prompting experience, said Shuker.

And a final tip from Ramp is to benchmark your spend per user. If it exceeds the median, find out why.

In other words, find out where you’re at and fast, before you become the headline.

Reading the room…

Here are the questions your board might be asking about your AI spend.

  • Indecent exposure: How large could our AI token spend become under plausible high-usage scenarios? What are the hard guardrails we have if any?

  • Cost capture: Do we know what we are spending in total today? How do we know we don’t have a shock around the corner? Do we have any shadow expense buried in supplier contracts or on employee expense claims?

  • Value realization: Do we have any actual measurable productivity benefit yet? Do we understand how that benefit is being realized or reinvested? Do we have the FP&A guardrails to protect it?

  • Model strategy: Are we routing work to the cheapest model capable of delivering acceptable quality?

  • Vendor dependence: How exposed are we to pricing changes, usage opacity, or lock-in from major providers?

Boardroom Brief is presented by The Secret CFO Network

Last week’s Playbook kicked off our deep dive into the financial history of Manchester United.

If you found this helpful, please forward it to your fellow finance leaders (and maybe even your Board). If this was forwarded to you, make sure you receive the next edition by subscribing here.