Microsoft Pulls Claude Code Licenses and Exposes the Real Problem with AI Costs: Technology Can End Up Costing More Than Employees
Microsoft is at the center of a debate that promises to shake up the tech industry in the years ahead. The company canceled most direct Claude Code licenses for its engineers, migrating everything to GitHub Copilot CLI, after just six months of making the tool available internally.
The reason? Adoption happened too fast, and the bill arrived sooner than expected.
According to The Verge, this move does not affect Microsoft’s deal with Anthropic through Foundry, which includes an investment of up to 5 billion dollars in Claude’s creator and a 30 billion dollar commitment from Anthropic to purchase computing capacity on Azure. The cut was specifically aimed at internal Claude Code licenses, which thousands of developers, project managers, designers, and other company employees were using daily for coding tasks.
This move raised an important red flag: running artificial intelligence at scale inside a company can end up costing more than paying human employees’ salaries.
And Microsoft is not alone in this.
Uber, for example, burned through its entire 2026 AI coding tools budget in just four months. The company’s CTO, Praveen Neppalli Naga, confirmed the situation to The Information in April, revealing that the company had actively encouraged the use of these tools through internal leaderboards that ranked teams by their volume of AI consumption.
Bryan Catanzaro, vice president of applied deep learning at Nvidia, was blunt in a recent interview with Axios:
For my team, the cost of compute far exceeds the cost of the employees.
That statement pretty much sums up the paradox that major tech companies are starting to face head-on — and it calls into question some of the promises made about the financial return of generative AI. 🤔
What Is Behind This Hefty Bill
When companies started distributing artificial intelligence tools to their teams, the idea was straightforward: boost productivity, cut time spent on repetitive tasks, and ultimately save money. The problem is that equation works great on paper, but in practice, the consumption of computing resources grows in ways that few managers anticipated. Every query sent to a language model, every line of code generated by an assistant, every automated analysis carries a real infrastructure cost. And when you multiply that by hundreds or thousands of employees using these tools all day long, the final number can be staggering.
In Microsoft’s case, Anthropic’s Claude Code was being used directly by engineers for programming and automation tasks. The tool is powerful, but the pay-per-use pricing model means that the more the teams work, the bigger the bill grows. Six months was all it took for the company to realize that usage volume was far beyond what was originally projected, and that continuing with direct licenses would represent spending that was hard to justify without a deeper analysis of return on investment. The solution was to consolidate everything under GitHub Copilot CLI, a platform that Microsoft already controls and can manage with more cost predictability.
This kind of move reveals something important about how the market is still learning to handle AI at enterprise scale. It is not enough to offer the tool and hope the results show up. You need to monitor consumption closely, create clear usage policies, and understand that the operational cost of advanced language models is a variable that can spiral out of control very quickly if there is no proper governance in place. The technology is incredible, but it needs to fit within a realistic financial plan.
Cheaper Tokens, Higher Bills: The Emerging Paradox
Microsoft and Uber are not isolated cases. The behavior of encouraging massive AI use has spread across the biggest tech companies in the world. At Meta, an employee created a dashboard called Claudeonomics — a nod to Anthropic’s Claude model — to track which workers were consuming the most AI. Amazon, meanwhile, started encouraging its employees to engage in so-called tokenmaxxing, which basically means using as many AI tokens as possible — the basic building blocks of computation for these models.
This widespread behavior of maximizing usage ends up creating a structural problem. With token-based pricing, work gets more expensive as usage increases and models become more sophisticated. Goldman Sachs recently published a projection showing that agentic AI could drive a 24x increase in token consumption by 2030, as consumers and businesses adopt AI agents at scale. The estimate is a staggering 120 quadrillion tokens per month. As businesses turn to AI agents to boost productivity, aggregate costs can spike sharply, even if the individual price of each token drops.
And yes, the unit cost of tokens is expected to decrease significantly. A recent Gartner report noted that by 2030, inference on a trillion-parameter language model — essentially, an extremely sophisticated AI model — will cost AI companies nearly 90% less than it did in 2025. Sounds like good news, right? Not quite.
Gartner itself warned that cheaper tokens will not translate into cheaper enterprise AI. And the reasons are pretty clear:
- Agentic models require far more tokens per task than standard models
- The increase in consumption can outpace the drop in unit costs
- AI providers will not fully pass cost reductions on to customers
Will Sommer, a senior analyst at Gartner, issued a pointed warning about this scenario: Product leaders should not confuse commodity token deflation with the democratization of frontier reasoning.
In plain terms: the fact that each individual token gets cheaper does not mean that operating advanced AI systems will cost less. In reality, it could cost significantly more, because the most advanced models simply consume far more tokens to accomplish their tasks.
When AI Costs More Than People
Bryan Catanzaro’s statement from Nvidia was not rhetorical exaggeration. It reflects a reality that is becoming increasingly common at companies operating on the frontier of artificial intelligence. R&D teams working with large models, training experiments, and large-scale inference already deal with compute budgets that easily surpass payroll. This happens because the computing power needed to run these systems is not cheap, especially when you are talking about cutting-edge GPUs and cloud infrastructure running continuously for days or weeks on end.
Uber’s case is perhaps the most striking example of this new reality. The company had a budget planned for all of 2026, earmarked for internal teams’ use of AI tools, and that budget was consumed in just four months. Not because the tools were misused or because there was obvious waste, but simply because the push to adopt them worked too well. Teams embraced the tools, became reliant on them in their daily workflow, and consumption scaled in ways that financial forecasting models did not accurately capture. When adoption is a success, the bill can become a problem — and that is a paradox few companies were prepared to face.
All of this puts new pressure on technology leaders and CFOs who need to justify these investments to their boards. The promise of generative AI has always been that it would pay for itself through productivity gains, but calculating that return with precision is still a massive challenge. How many work hours were saved? How many bugs were avoided? How many decisions were made faster? These are valid questions, but the answers rarely come with the clarity needed to balance a spending spreadsheet that keeps growing month over month with infrastructure and licensing costs.
The Vision of 100 AI Agents per Employee Will Come with a Price Tag
Jensen Huang, CEO of Nvidia, recently stated that he envisions a future where 100 AI agents will work alongside each employee within his own company. This vision is part of a broader wave of CEOs at major corporations promoting an agentic future, where digital workers operate autonomously across different areas of organizations.
The idea is exciting, no doubt. But the reports from Microsoft and Uber suggest that if token consumption grows faster than unit costs fall, this future could arrive with a bill much heavier than executives imagine today. The complexity of the tasks AI agents need to solve requires long reasoning chains, multiple processing steps, and enormous volumes of tokens for each completed operation.
Putting 100 agents to work for each employee is not just a question of technical capability. It is, first and foremost, an economic question that needs to add up at the end of the month. And based on what the data is showing right now, that math still does not work for most companies that have tried to scale AI usage internally. 📊
What Changes Going Forward
The current landscape does not mean companies are going to give up on artificial intelligence. Quite the opposite. But it does signal that the phase of unrestricted, enthusiastic distribution is coming to an end, giving way to a more strategic and deliberate approach to usage. Organizations are starting to treat AI consumption with the same level of rigor they apply to any other corporate resource — setting usage limits, categorizing use cases by value generated, and prioritizing tools that deliver measurable results instead of simply adopting everything that hits the market.
For employees who use these tools day to day, this could mean concrete changes in how they access AI at work. Stricter usage policies, approvals for certain types of queries, or even replacing more expensive tools with more affordable internally controlled alternatives are all moves that should become more common over the coming months. Microsoft already set a clear example by migrating its engineers from Claude Code to GitHub Copilot CLI — a decision that blends cost control with platform consolidation within the company’s own ecosystem.
Another relevant point is that neither Anthropic nor Microsoft publicly commented on the details of this decision when contacted by Fortune, which suggests the topic is sensitive internally and that the companies are still calibrating how to communicate these changes without undermining the innovation narrative that underpins their market strategies.
Pricing Needs to Evolve
The big question that remains is whether pricing models for AI tools will evolve to match this new reality. Today, most charges are based on usage volume, which works fine for occasional use but becomes problematic when usage is intense and continuous. There are possible paths that could make this equation more sustainable:
- Outcome-based models, where charges are tied to the value delivered rather than the volume consumed
- Subscriptions with consumption caps that offer budget predictability
- Enterprise packages designed for large-scale use with progressive discounts
- Proprietary internal tools that reduce dependence on external vendors
These models could be the way to make the equation more balanced, both for the tech companies selling the tools and for the organizations that need to justify every line item on their innovation budget.
Market Maturity Runs Through This
What we are witnessing right now is actually a sign that the enterprise AI market is maturing. Every disruptive technology goes through this cycle: initial enthusiasm, accelerated adoption, a reality check on costs, and finally, stabilization with smarter and more sustainable usage models. Generative artificial intelligence will be no different.
The reports about Microsoft and Uber are not stories about AI failing. They are stories about companies learning, in real time, how to calibrate expectations against operational reality. And that calibration is essential for the technology to actually deliver the value it promises in the long run. The companies that manage to find this balance between innovation and financial responsibility will likely be the ones leading the next phase of this technological transformation. 💡
In the meantime, the message for the market is clear: AI is powerful, transformative, and full of potential — but it is not free. And pretending that cost does not matter is the fastest recipe for turning a technological revolution into a problem on the balance sheet.
