Back to News
Market Impact: 0.25

Former OpenAI CTO does what Altman won't: releases a frontier AI model that's actually open

NVDA
TSTS
Artificial IntelligenceTechnology & InnovationCompany FundamentalsPrivate Markets & Venture

Thinking Machines Lab released “Inkling,” a 975B-parameter open-weights reasoning LLM under an Apache 2.0 license, available via its Tinker platform today and on Hugging Face for local download. The model targets large-context use cases (up to 1M tokens) and is offered in a NVFP4 quantized variant for lighter GPU needs, positioning it as the largest American open-weights model to date. Management claims competitive performance versus frontier closed models but acknowledges benchmarks show trailing Anthropic’s Claude and OpenAI’s GPT, while highlighting efficiency improvements via reinforcement-learning “thinking tokens.”

Analysis

This is incrementally bullish for NVDA not because one model matters, but because it reinforces the economic regime where frontier capability requires dense GPU clusters, high-bandwidth memory, and networking even when the model is freely downloadable. The key second-order effect is that permissive open weights tend to expand experimentation and downstream inference volume faster than they reduce hardware demand; if developers can self-host and fine-tune, they still need rented or owned accelerators to do it.

The bigger pressure is on proprietary model monetization and API intermediaries. A credible open model narrows the gap between "best" and "good enough," which can compress pricing power for closed-model vendors over the next 1-3 months as enterprise evals shift toward cost/performance and control. That said, the most likely outcome is not a collapse in model spend but a mix shift from software subscription economics toward infrastructure economics, which is structurally friendlier to semis and GPU cloud platforms than to pure model layers.

Contrarian view: the market may underappreciate how much quantization and MoE architecture can lower effective GPU demand per unit of capability. If this approach becomes the default, capability growth may not translate one-for-one into transistor demand, which is the main bear case for NVDA over 6-18 months. Falsifiers are straightforward: if hyperscaler capex commentary, NVDA data-center backlog, or third-party inference utilization do not improve after these launches, the hardware bull thesis is probably overstated.