๐ค AI ้่ง
๐ ๆ็ซ ๅ ๆฐๆฎ
- ๅๅธๆถ้ด
- 2026-06-22
- ็ฑปๅ
- posts
- ๆ ็ญพ
- AI, Token Economics, Big Tech, Cost Crisis, Deep Analysis
1. Bill Shock: Silicon Valley’s Fatal Four Months Link to heading
Midway through 2026, the AI industry is undergoing an unexpected awakening. Four things happened almost simultaneously:
Microsoft urgently canceled Claude Code licenses for most employees in May, forcing tens of thousands of engineers back onto GitHub Copilot CLI. The pilot program, launched with fanfare in December 2025, collapsed after just six months due to runaway bills.
Uber’s CTO admitted internally that the company’s $3.4 billion AI budget for 2026 was already exhausted by April. After deploying Claude Code to 5,000 engineers, monthly adoption soared to 95%, with per-engineer token costs running $500 to $2,000 per month.
Amazon shut down its internal AI usage leaderboard, KiroRank, in May. Employees had been gaming the system by having AI agents perform meaningless tasks en masse to boost their rankings, artificially inflating compute consumption โ giving rise to a new term: Tokenmaxxing.
An unnamed tech giant (industry speculation points to a Microsoft or Meta subsidiary) racked up a $500 million monthly bill after forgetting to set usage caps on employee Claude licenses โ burning over 100 million RMB per day.
And this wasn’t isolated. Meta, Amazon, and Uber had all enthusiastically tied token consumption to performance reviews: whoever burned the most was the most “innovative.” The result? Employees turned it into a creative writing contest in cost inflation.
2. For Every $1 in Tokens, Only 18 Cents Creates Value Link to heading
Developer productivity platform Entelligence.AI aggregated data from 2,444 companies. The findings are sobering:
For every $1 spent on AI tokens, only 18 cents generates actual user-facing value.
44 cents goes to fixing bugs introduced by AI itself. 27 cents flows to rework. 11 cents is consumed by review friction.
In other words, over 80% of token expenditure is paying for AI’s immaturity.
And that’s before accounting for internal management black holes:
- Li Cong, founder of Onion Group, stated bluntly: “Many employees are using company tokens to slack off, even moonlight. During the day at work, they take on external dev gigs, design jobs, and operations tasks using company-issued model quotas.”
- Amazon’s Tokenmaxxing not only wasted compute resources โ the finance department received bills where “the data was already three days stale.”
- One team running a Multi-Agent experiment blew past a million RMB in token costs in a single night.
When your KPI becomes “how many tokens can you burn,” the most rational employee response is unlimited token consumption to game the score. This is Goodhart’s Law in perfect execution: “When a measure becomes a target, it ceases to be a good measure.”
3. Why Tokens Went from Dirt Cheap to Luxury Goods Link to heading
The key number: token prices have risen roughly 65% since February 2026. GPT-5.5 doubled its pricing. Gemini raised rates 3x in some scenarios. Claude’s API costs have surged.
Three layers are driving this:
1. Structural Supply Constraints
- HBM (High Bandwidth Memory): Samsung, SK Hynix, and Micron control 95% of global capacity. Their expansion cycle takes 24-36 months. HBM prices rose over 50% since late 2025.
- CoWoS advanced packaging: even after TSMC doubled capacity, 2026 orders remain backlogged through year-end.
- An 8-GPU NVIDIA B300 server went from under 4 million RMB to roughly 7 million RMB, and “sells out the moment it arrives.”
2. Explosive Demand Growth
- Global weekly token consumption surged from 2.1T to 24.5T in one year, with 2026 year-over-year growth at 280%.
- Google Gemini’s monthly token volume jumped from 480 trillion to 3,200 trillion โ over 6x growth.
- Epoch AI report: global Blackwell chip compute capacity grows roughly 3.4x annually, while token demand grows roughly 10x annually. 3.4 vs. 10 โ the gap widens every year.
3. The Agent Multiplier
When AI shifts from simple Q&A to autonomous execution, token consumption jumps from hundreds to hundreds of thousands, even millions, per task.
Take a composite “book flight + hotel + rental car” task: user input accounts for under 1%, model reasoning about 5-10%, and tool calls (API interactions) about 85-90%. The final output is under 5%.
This means optimizing model inference alone has sharply limited cost-saving potential โ the real consumption comes from agents repeatedly interacting with external environments. As Uber’s CTO put it: when AI evolves from a chat box into “autonomously refactoring your entire project, running loops wildly in the background,” it burns through tokens “like a printing press with no off switch.”
4. NYU Professor’s Warning: This Could Hurt More Than Dot-Com Link to heading
Aswath Damodaran, NYU finance professor and the “Dean of Valuation,” recently issued an unsettling warning:
“The AI boom has one key difference from the dot-com era: it requires massive investment in physical infrastructure like data centers, and a significant portion is financed through debt. When the correction comes, the losses won’t stay confined to shareholders โ they could spread throughout society.”
He further notes that the dot-com era involved almost no major capital expenditure. Today’s AI boom, by contrast, consists of real physical asset investments in data centers and chip procurement, largely financed through private debt rather than traditional banking. If defaults and distress follow, “the pain won’t stop at the corporate level.”
Damodaran also questions whether AI’s business model can ever achieve the scale economics investors hope for: “AI is different from traditional software. More users don’t necessarily mean unit costs naturally trend toward zero.”
5. From Burning Tokens to Managing Tokens Link to heading
The good news: the industry isn’t sitting still.
Companies are self-correcting:
- Uber has imposed monthly spending caps on engineer AI tool usage.
- Coinbase set up weekly budgets of $500-$5,000 per employee by seniority level.
- Amazon shifted its metric from token consumption to “standardized deployment” โ measuring actual AI-assisted code delivered.
- Tencent restructured its token allocation, ending flat-rate distribution in favor of department-managed dynamic assignment.
A new role is emerging: FinAI.
Just as cloud computing gave rise to FinOps for cost management, DeepAI CEO Cui Wei predicts a dedicated AI cost management role will emerge. Which departments need top-tier models? Which tasks are fine with 70-point models? What scenarios truly justify burning Claude credits, and where is it unnecessary? These require granular AI financial discipline.
Demand for “budget models” is surging:
Coinbase and others have started routing basic workloads to lighter Chinese models. Code agent startup Command Code gained 10,000 new customers in 30 days purely from demand for cheaper alternatives. As one Silicon Valley engineer put it: “Using a top-tier AI model to write your weekly status report is like taking a Ferrari to buy groceries.”
The Linux Foundation is stepping in:
In July, the Linux Foundation will formally launch the “Tokenomics Foundation,” backed by IBM, Oracle, and JPMorgan Chase. Its mission: establish new metrics like “cost per unit of intelligence” and “tokens per watt,” bringing financial discipline to the still-wild west of AI spending.
6. Conclusion: This Isn’t AI’s Decline โ It’s Its Coming-of-Age Link to heading
Some say the Great Token Retreat signals an impending AI bubble burst. But I think a more accurate description is: this isn’t “AI is failing” โ this is “AI is starting to do the math.”
The past two years ran on FOMO: doesn’t matter if AI works, just get on board. All costs were masked by model company subsidies. Entering 2026, subsidies are retreating, bills are surfacing, and enterprises are demanding verifiable ROI from AI.
This is precisely what industry maturity looks like.
The real value isn’t in “how many tokens can you burn,” but in “how much value does each token create.” The industry consensus is shifting from “who burns harder” to “who manages better.”
Token scarcity isn’t a technology problem โ it’s an economics problem. It reminds us all: compute may be vast, but it isn’t infinite. Efficiency may be high, but it isn’t free. Innovation may be good, but it isn’t costless.
Next time you open a chat window and burn a few hundred tokens on a simple question, spare a thought for the agents looping tens of thousands of reasoning cycles in the background, the automated workflows repeatedly calling external tools, the convoluted inference chains chasing down a single bug โ behind every token is a real bill coming due.
And that may be where true AI commercialization finally begins.
Sources: Entelligence.AI Enterprise Token Efficiency Report; Sina Finance / Pencil News “Token Great Retreat” deep dive; TMTPost “Giants Can No Longer Afford to Burn Tokens”; The Decoder / Livemint Damodaran interview; Business Insider / Financial Times Amazon Tokenmaxxing coverage; Wallstreetcn AI supply chain analysis; Zhitong Finance Token cost crisis synthesis