Back to News
Market Impact: 0.25

McKinsey: Cheaper AI models, bigger AI bills

Source: Fortune

Artificial IntelligenceTechnology & InnovationCompany FundamentalsCorporate Guidance & Outlook

McKinsey says GPT-4-class model inference costs have fallen from roughly $60 per million output tokens in early 2023 to below $1 in some cases, yet enterprise AI spending is rising as agentic systems consume substantially more reasoning capacity. Agent-task costs can vary by as much as 30x between runs, making workflow design, model routing, caching, usage visibility and vendor sourcing central to realizing ROI. McKinsey argues enterprises should optimize AI spending for task-level value and verification time rather than pursue indiscriminate cost cuts.

Analysis

The key investable mechanism is a Jevons-paradox outcome: falling inference unit costs expand the addressable task set faster than they lower aggregate spend. That favors hyperscalers with vertically integrated distribution and capacity—MSFT, AMZN and GOOGL—because enterprise buyers will increasingly optimize at the workload level while remaining dependent on their identity, data, security and cloud-control planes. The near-term revenue effect is usage-led rather than seat-led, supporting cloud growth through the next 1-3 quarters even if API pricing continues to fall.

The less obvious beneficiary is the AI governance/observability layer. As enterprises discover that agent workloads have highly dispersed cost and reliability profiles, tooling that attributes spend, traces failures and enforces guardrails becomes a budget prerequisite rather than discretionary developer tooling; DDOG and NOW are plausible public proxies. ServiceNow is especially positioned for 6-18 month upside if agent deployment shifts from isolated copilots toward redesigned approval, service and back-office workflows, where it can monetize orchestration and retain human-in-the-loop controls.

Consensus may over-focus on model-price deflation as margin-negative for AI vendors. Lower pricing is negative only if workload growth decelerates below efficiency gains; the more immediate risk is that enterprises deploy agents without measurable labor or revenue capture, producing a 2026 procurement pause after initial experimentation. Falsify the infrastructure thesis if hyperscalers report AI workload growth but cloud consumption growth and remaining-performance-obligation trends fail to accelerate over two earnings cycles, or if enterprise AI budgets shift materially from consumption contracts to capped-seat licenses.

AllMind Terminal

AI-powered research, real-time alerts, and portfolio analytics for institutional investors.

Request Trial

Market Sentiment

Overall Sentiment

mildly positive

Sentiment Score

0.25

Key Decisions for Investors

  • Maintain a 3-6 month overweight in AMZN versus enterprise-software peers with limited usage exposure: AWS captures incremental agent inference and workflow data movement, while retail cash flow limits capex-balance-sheet risk. Reassess if AWS growth fails to reaccelerate or AI capex drives a material deterioration in free-cash-flow conversion.
  • Initiate a 6-12 month pair trade long NOW / short equal-dollar basket of mature seat-based SaaS exposure via IGV: workflow redesign and governance should be monetizable even when customers rationalize redundant licenses. Target 15-20% relative upside; stop if NOW's subscription cRPO growth decelerates by more than 300 bps or AI products show no attach-rate disclosure by the next two earnings reports.
  • Keep DDOG on an earnings-watch list rather than initiate immediately: a long is warranted only if management quantifies AI/LLM observability contribution or reports sustained usage acceleration alongside stable net retention. The missing data is whether AI tracing converts into material paid consumption rather than free experimentation.
  • Avoid treating falling model costs as a standalone short catalyst for MSFT, AMZN or GOOGL over the next quarter; use any multiple-driven pullback to add exposure. The adverse scenario is capacity oversupply, visible through declining cloud revenue growth despite continued capex escalation, which would warrant reducing the basket.

More News

From AllMind Research

Browse all research