Blog · Strategy

Microsoft and Meta cut their Claude bill. Anthropic’s enterprise share still went up.

Both of those are true, and the reason has more to do with who is shipping a model than with the price of a token. What the gap between them means for your AI budget.

a thirdof Microsoft’s internal Claude spend was cut
40%of enterprise LLM API spend went to Anthropic in 2025, up from 12% in 2023
95%of enterprise GenAI pilots produced no measurable return

Sources: The Information, via Yahoo Finance; Menlo Ventures, 2025: The State of Generative AI in the Enterprise; MIT Project NANDA, The GenAI Divide.

In short

  • Microsoft cut internal Claude spend by over a third. Meta cut Claude Code use roughly in half. Neither stopped the work.
  • In the same year, Anthropic went to 40% of enterprise LLM API spend and 54% of coding spend. The headline and the market point in opposite directions.
  • Both companies ship their own models now. Moving your engineers onto your own model is a product decision before it is a cost decision.
  • Routing cheap work to cheap models is sound, and it is not where enterprise AI is failing. 95% of pilots return nothing because nobody measured the process first.

A link went round our Slack on a Monday: Meta and Microsoft are scaling back internal use of Claude. The thread that followed was better than the article, and it came down to a disagreement about what the story was even about.

The reporting itself is solid. The Information found that Microsoft cut internal Claude spending by more than a third and is cancelling most Claude Code licences across its Experiences and Devices division, pointing engineers working on Windows, Microsoft 365, Outlook, Teams and Surface at GitHub Copilot CLI instead. Meta cut Claude Code usage roughly in half. The reason given in both cases is cost. The tools got popular, and the token bill followed.

The reading most people took from it is the one I keep hearing inside other companies too: the spend is hard to justify when nobody can point at what it bought. That is a real problem and I will come back to it, because it is the part of this story that actually applies to everyone.

But cost is the reported reason. It is rarely the whole one.

Two of Claude’s biggest internal users cut back in the year enterprise spend on Claude went up

If Claude were losing on the merits you would expect the market to move with it. It moved the other way. Menlo Ventures put Anthropic at 40% of enterprise LLM API spend in 2025, up from 24% in 2024 and 12% in 2023.

OpenAI went the other direction over the same period, from 50% in 2023 to 27%.

Google went from 7% to 21%.

Foundation model API spend reached $12.5 billion for the year. In coding specifically, the category this story is actually about, Anthropic held 54% against OpenAI’s 21%.

So the two facts sit side by side: a handful of very large customers reduced internal usage, and the category kept consolidating toward the same vendor. Both are true. Treating the first as a verdict on the second is the mistake.

When a company builds its own model, its own engineers are the training set

Here is what I think is actually happening, and it is not mainly about the bill.

Meta is no longer only a customer. Meta Superintelligence Labs introduced Muse Spark in April 2026, shipped Muse Code as a terminal coding agent alongside 1.2, and released Muse Spark 1.3 in September, priced at $1.25 per million input tokens and $4.25 per million output tokens. A company shipping a coding model needs its own engineers on it, for two reasons that have nothing to do with the invoice. The first is data: your engineers at work are the highest-quality trace data you will ever get for a coding model, and you cannot collect it while they are using somebody else’s. The second is simpler. You cannot credibly sell a coding model that your own engineers decline to use.

There is a contractual edge to this too. Anthropic’s commercial terms say a customer may not access the services to build a competing product or service, including to train competing AI models, and every major provider has a version of that clause. None of this means anyone has broken it. It means that once you are shipping a rival model, the tidiest position is the one where the question never has to be asked, and moving your engineers off is how you get there.

So this is a product decision before it is a procurement one.

Microsoft is in a more awkward spot than Meta

Microsoft is a genuinely mixed case. Its customer-facing AI products lean heavily on OpenAI, while internally a lot of its developers reached for Anthropic and kept reaching. The reporting says Claude Code had become, in the words of one account, a little too popular. That is not a procurement failure. That is a preference showing up in a cost line.

Microsoft would like a home-grown model carrying that load, and so far nothing it has built has closed the gap. So what reads as a retreat is better described as a reallocation: the same engineering work moved to a cheaper tool, not stopped. My guess is that the enterprise-scale contracts get trimmed first and individual seats stay, because for most organisations a per-head subscription is cheaper than a negotiated deal at volume.

Cutting a bill is not the same as cutting the work. Those get reported as one thing and they are not.

Most of your AI work does not need your most expensive model

The part of this that generalises is the routing, and here the economics are genuinely on your side. Stanford HAI’s 2025 AI Index found the cost of querying a model at roughly GPT-3.5-level performance on MMLU fell from $20 per million tokens in November 2022 to $0.07 by October 2024. That is about 280x in 18 months, and the capability floor has kept rising since.

What that buys you is a Pareto curve to ride. Summarising, classifying, extracting, tagging and first-draft writing clear the quality bar on cheap models now, and they are the bulk of what a non-engineering organisation actually does with AI. Engineering is the exception and will stay the exception for a while, even with strong mid-tier options like GPT-6 Sol in the mix. If you are paying frontier prices for work a cheap model handles, that is worth fixing, and it is the one lesson from the Microsoft story that transfers cleanly.

It is worth being concrete about what sits where, because “use a cheaper model” is not a plan. This is the split I work to when I look at a questionnaire workflow:

Which model tier each step of a questionnaire workflow actually needs
The workUsual defaultWhat it actually needs
Reading a questionnaire into structured questionsFrontier modelCheap. It is parsing, not judgment.
Drafting a first answer from approved sourcesFrontier modelMid-tier. The constraint is the source, not the reasoning.
Deciding an answer needs a humanNobody decidesA classifier and a threshold you set
Reconciling two approved answers that disagreeFrontier modelFrontier. This is the judgment call.
Production codeFrontier modelFrontier, still

Two of those five need the expensive model. Getting that ratio right is most of what the Microsoft story is actually about, and it is the part you can act on this quarter.

Routing is not free, though. Somebody has to own the decision about which work goes where, and be willing to be wrong in public when a cheaper model ships a worse answer to a customer. Teams that skip that ownership step end up with a routing layer nobody trusts and a quiet drift back to the expensive default.

The spend is not what is failing. The measurement is.

Now back to the point about justifying the spend, because it is the one that matters most.

MIT’s Project NANDA reviewed more than 300 public initiatives, interviewed 52 organisations and surveyed 153 senior leaders for a report called The GenAI Divide. It found that 95% of enterprise generative AI pilots produced no measurable P&L return. The authors put the gap down to brittle workflows, weak contextual learning and poor fit with how people actually work, not to model quality and not to price.

Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.

Read those next to each other and the cost story inverts. A pilot that returns nothing is expensive at any token price. A deployment that takes four hours off a process somebody already reports on every month is cheap even at frontier rates. The question was never cost per token. It was whether anyone could tell you what the thing was worth, and in 95% of cases nobody could, because nothing was being measured before the pilot started.

It is also why Tribble Respond points at questionnaires and not at knowledge in general. An RFP has a clock on it, a queue behind it, and a reviewer who already knows what share of answers they had to rewrite. The baseline is sitting there whether or not anybody has called it one, which means the before and after are arguable in a way that “the team feels faster” never is.

The shadow AI economy is the actual forecast

The same MIT report found something that gets quoted less and predicts more: staff at over 90% of companies use personal AI tools for work, while only 40% of companies provide official access.

That is the end state I would bet on for organisations where the product is not built by engineers. Not a large enterprise agreement. A budget line, a sensible list of approved tools, and people picking what suits the task. It is cheaper than a negotiated contract and it is already how most of these companies operate, whether or not anyone has written it down.

It also has a cost that does not show up on the invoice. When the work happens in personal accounts, the answers people depend on live in someone’s chat history. There is no source on them, no approval, no record of who decided what, and no way to correct an answer once it is wrong in forty places. That is fine for drafting an email. It is not fine for the answer you send a regulator, a security reviewer or a customer.

That gap is what the Brain is for. Every answer Tribble gives carries where it came from, who owns it, and who was allowed to see it, so “who approved this, and when” is a question with an answer. A personal chat account cannot do that, and it is not trying to.

Which is where this stops being a story about model vendors.

The teams getting a return picked work that already had a number on it

The deployments I see produce a measurable result are not the ambitious ones. They are the ones pointed at a process the business was already counting: how long an RFP takes, how many security questionnaires are in the queue, how much of a response needed a specialist to rewrite it. The number existed before the AI did, which is the only reason anyone can tell whether it moved.

That is the work Tribble Respond is built for. A questionnaire lands, Tribble reads the questions out of whatever file they arrived in, drafts from approved sources only, shows the source on every answer, and flags the ones that need a human instead of guessing. It scores how sure it is, and anything below the line you set goes to a person. Customers started calling the governed layer underneath it the Brain before we did.

“I completed 90% of a 200-question RFP in just an hour, and it handles one-off technical questions from sales with incredible accuracy. Now I can focus more time on strategic customer engagement instead of searching for information.”

Shaundra Toy · Sales Engineer, Clari

That is the shape of number I mean. Not “faster”, but 90% of a known workload in a known time, against a process Clari was already measuring.

The number I find more interesting is the other one from the same deployment: 10-20% of security responses still needed a specialist. That is the figure that makes the first one trustworthy. A system that cannot tell you which tenth needs a human has not saved you the work, it has moved the risk somewhere you cannot see it.

Then the second-order thing happens, which is the part I did not expect when I started building this. The repeated questions turn into a signal of their own: what buyers keep asking, where an approved answer has gone stale, what the company is not yet ready to answer. That is what Tribblytics reads. Clari’s team went from treating responses as administrative work to reading them for product gaps. That only works because the answers were governed to begin with, which is the argument of this whole post one layer down.

What Tribble does not do is make the measurement problem go away. If nobody can tell you today how long your RFP process takes or what share of answers get rewritten, that is the first week of work, and it is yours. A tool cannot retrofit a baseline you never had.

The companies in the headline are not telling you Claude stopped working. They are telling you they would rather own the model. You probably would not, and that is the one part of their strategy worth copying: know exactly what the spend is for.

FAQ

Frequently asked questions

What people ask about the Claude cutbacks and enterprise AI spend.

Did Microsoft and Meta stop using Claude?

No. The Information reported that Microsoft cut internal Claude spending by more than a third and cancelled most Claude Code licences in its Experiences and Devices division, and that Meta cut Claude Code usage roughly in half. Both are reductions in internal usage, not removals, and both companies moved that work to other tools rather than stopping it.

Is enterprise spending on Anthropic falling?

The opposite, at the market level. Menlo Ventures put Anthropic at 40% of enterprise LLM API spend in 2025, up from 24% in 2024 and 12% in 2023, against OpenAI at 27% and Google at 21%. In coding specifically Anthropic was at 54% against OpenAI’s 21%. Two large internal users cutting back is compatible with the overall market growing.

Should we switch to a cheaper model to cut AI costs?

For some of the work, yes. Stanford HAI’s AI Index found the cost of querying a model at GPT-3.5-level performance on MMLU fell from $20 per million tokens in November 2022 to $0.07 by October 2024, roughly 280x in 18 months. Summarising, classifying, extracting and first-draft writing usually clear the bar on cheaper models. Engineering work is the exception, and routing has a real cost: somebody has to own which work goes where and be willing to be wrong about it.

Why do most enterprise AI projects show no return?

Usually because nothing was being measured before the pilot started. MIT’s Project NANDA reviewed more than 300 public initiatives and found 95% of enterprise generative AI pilots produced no measurable P&L return, attributing the gap to brittle workflows and poor fit with daily operations rather than to model quality or price. Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027 for similar reasons.

What is shadow AI and why does it matter for budgets?

Shadow AI is employees using personal AI subscriptions for work. MIT’s Project NANDA found staff at over 90% of companies do this while only 40% of companies buy official access. It matters for budgets because it is often where the real productivity gain sits, and it matters for governance because the answers people rely on then live in personal chat histories, with no source, no approval and no audit trail.

Bring a real RFP and we will answer it live

Thirty minutes, nothing to prepare. Bring one of yours if it is handy. If not, we will use ours.

Book a demo