McKinsey: Cheaper AI models, bigger AI bills
Source: Fortune
McKinsey says GPT-4-class model inference costs have fallen from roughly $60 per million output tokens in early 2023 to below $1 in some cases, yet enterprise AI spending is rising as agentic systems consume substantially more reasoning capacity. Agent-task costs can vary by as much as 30x between runs, making workflow design, model routing, caching, usage visibility and vendor sourcing central to realizing ROI. McKinsey argues enterprises should optimize AI spending for task-level value and verification time rather than pursue indiscriminate cost cuts.
Analysis
The key investable mechanism is a Jevons-paradox outcome: falling inference unit costs expand the addressable task set faster than they lower aggregate spend. That favors hyperscalers with vertically integrated distribution and capacity—MSFT, AMZN and GOOGL—because enterprise buyers will increasingly optimize at the workload level while remaining dependent on their identity, data, security and cloud-control planes. The near-term revenue effect is usage-led rather than seat-led, supporting cloud growth through the next 1-3 quarters even if API pricing continues to fall.
The less obvious beneficiary is the AI governance/observability layer. As enterprises discover that agent workloads have highly dispersed cost and reliability profiles, tooling that attributes spend, traces failures and enforces guardrails becomes a budget prerequisite rather than discretionary developer tooling; DDOG and NOW are plausible public proxies. ServiceNow is especially positioned for 6-18 month upside if agent deployment shifts from isolated copilots toward redesigned approval, service and back-office workflows, where it can monetize orchestration and retain human-in-the-loop controls.
Consensus may over-focus on model-price deflation as margin-negative for AI vendors. Lower pricing is negative only if workload growth decelerates below efficiency gains; the more immediate risk is that enterprises deploy agents without measurable labor or revenue capture, producing a 2026 procurement pause after initial experimentation. Falsify the infrastructure thesis if hyperscalers report AI workload growth but cloud consumption growth and remaining-performance-obligation trends fail to accelerate over two earnings cycles, or if enterprise AI budgets shift materially from consumption contracts to capped-seat licenses.
AllMind Terminal
AI-powered research, real-time alerts, and portfolio analytics for institutional investors.
Request TrialMarket Sentiment
Overall Sentiment
mildly positive
Sentiment Score
0.25
Key Decisions for Investors
- Maintain a 3-6 month overweight in AMZN versus enterprise-software peers with limited usage exposure: AWS captures incremental agent inference and workflow data movement, while retail cash flow limits capex-balance-sheet risk. Reassess if AWS growth fails to reaccelerate or AI capex drives a material deterioration in free-cash-flow conversion.
- Initiate a 6-12 month pair trade long NOW / short equal-dollar basket of mature seat-based SaaS exposure via IGV: workflow redesign and governance should be monetizable even when customers rationalize redundant licenses. Target 15-20% relative upside; stop if NOW's subscription cRPO growth decelerates by more than 300 bps or AI products show no attach-rate disclosure by the next two earnings reports.
- Keep DDOG on an earnings-watch list rather than initiate immediately: a long is warranted only if management quantifies AI/LLM observability contribution or reports sustained usage acceleration alongside stable net retention. The missing data is whether AI tracing converts into material paid consumption rather than free experimentation.
- Avoid treating falling model costs as a standalone short catalyst for MSFT, AMZN or GOOGL over the next quarter; use any multiple-driven pullback to add exposure. The adverse scenario is capacity oversupply, visible through declining cloud revenue growth despite continued capex escalation, which would warrant reducing the basket.
More News
- China's AI chip blitz arms Xi with a message for Trump: 'You can't choke us off'
- OpenAI’s agent hacked Australia’s Medicare website—the latest rogue AI incident that the company didn’t know about for months
- Meta announces new lightweight virtual reality glasses to one-up Apple’s Vision Pro
- Markets are rapidly coming around to the reality that the Fed has a lot more work to do
- Chinese authorities reportedly in possession of F-35 components in Hong Kong
- SoftBank shares jump over 7% after $11.1 billion bond issuance to fund OpenAI bet