
AI inference is shifting from training to production, with Deloitte estimating inference will account for two-thirds of AI data center computing power this year. Broadcom reported AI revenue up 143% year over year to $10.8 billion in Q2 fiscal 2026 and guided to $16 billion next quarter, while AMD said server CPU revenue should rise 70% year over year as inference demand accelerates. Despite competition from custom chips and CPUs, Nvidia remains dominant with a reported 74% share of AI inference chips and $41 billion in inference processor sales in Q1 2026.
The key second-order effect is that inference does not create a clean winner-take-all CPU/ASIC transition; it expands the total addressable compute layer while shifting mix toward systems integration, networking, and software optimization. That favors the vendor best able to monetize the entire stack rather than just silicon, which explains why the market keeps rewarding the most vertically integrated platform even as unit economics improve for lower-compute alternatives.
The biggest misconception is that cheaper inference should mechanically compress the incumbent’s share. In practice, lower per-query cost tends to expand usage faster than it cannibalizes accelerator demand, especially as agentic workflows multiply the number of model calls per user action. That dynamic can keep high-end GPU demand sticky for months, not quarters, because orchestration, retrieval, and fallback paths still require fast general-purpose compute even when the primary workload is inference.
For AMD and AVGO, the opportunity is real but more duration-sensitive: both benefit if hyperscalers diversify supply, yet their upside depends on sustained capex discipline and successful customer qualification cycles. The risk is that inference customization becomes a scale game dominated by a few hyperscalers, which would cap merchant silicon share and push more economics inward to the cloud operators themselves. A slowing of AI capex growth, or a shift toward smaller on-device models, would hit the smaller inference beneficiaries first.
Contrarian take: the market may be underestimating how much price-performance improvement can extend the runway for Nvidia rather than erode it. If next-gen platforms materially cut inference cost, the likely near-term effect is higher throughput utilization and faster model deployment, not lower revenue. The setup argues for staying long the platform leader on pullbacks, while being more selective on the inference-adjacent names where growth is faster but moat durability is less proven.
AI-powered research, real-time alerts, and portfolio analytics for institutional investors.
Request DemoOverall Sentiment
mildly positive
Sentiment Score
0.35
Ticker Sentiment