Back to News
Market Impact: 0.3

AI model costs are pushing startups towards cheaper open weigh

Source: The Next Web

Artificial IntelligenceCompany FundamentalsTechnology & Innovation

Harvey's gross margin collapsed from roughly 50% at the start of the year to negative 50% by June as usage of its AI agents surged, highlighting severe inference-cost pressure. Margins returned to positive territory only after Harvey launched a proprietary model built on Moonshot's Kimi K3. Abridge, Decagon, Ramp and Rogo are pursuing similar in-house model strategies to improve AI-unit economics.

Analysis

The relevant signal is not demand weakness but a potential reset in vertical-AI unit economics: application vendors that rely on third-party frontier-model APIs can see revenue scale faster than gross profit when agentic workflows increase token intensity. The firms able to route workloads to smaller, specialized, or self-hosted models should regain margin fastest; those without proprietary data, workflow lock-in, or model-optimization talent risk being forced into price competition before their revenue multiples are supported by durable cash generation. This is structurally negative for the weakest cohort of high-multiple private AI software companies, even if reported ARR remains strong over the next 1-3 quarters.

Public-market read-through is indirect. Hyperscalers (MSFT, GOOGL, AMZN) benefit from greater inference volume and enterprise deployment, but application-layer optimization may shift spend away from premium closed-model APIs toward lower-cost open-weight or customized models, limiting the assumption that every agentic workflow produces frontier-model pricing power. For 6-18 months, the likely winners are incumbent workflow vendors such as NOW and CRM if they can embed agents while preserving seat pricing and avoid usage-linked COGS leakage; investors should focus on AI gross-margin disclosure, inference-cost trends, and whether AI features are monetized separately rather than bundled.

There is no direct listed-equity trade in the named Ramp business: it is privately held, while Nasdaq-listed LiveRamp (RAMP) is a separate data-connectivity company. Treat any apparent RAMP linkage as ticker contamination rather than an investable catalyst. The contrarian view is that model-cost compression ultimately expands adoption enough to offset lower unit pricing, but that outcome requires AI revenue to grow materially faster than inference expense for at least two reporting cycles.

AllMind Terminal

AI-powered research, real-time alerts, and portfolio analytics for institutional investors.

Request Trial

Market Sentiment

Overall Sentiment

mixed

Sentiment Score

-0.10

Ticker Sentiment

RAMP0.15

Key Decisions for Investors

  • No directional trade in RAMP from this item; verify issuer identity before acting, as the article's Ramp is private and unrelated to LiveRamp (RAMP).
  • Establish a 1-3 month watchlist on NOW and CRM: consider longs only after earnings show AI-related attach revenue or stable/improving subscription gross margin despite increased agent usage. Falsifier: AI adoption rises while gross margin falls by more than 100 bps or management cannot quantify monetization.
  • Use a relative-value screen across public AI application software: underweight names with heavy usage-based AI features, limited proprietary data, and declining gross margin versus NOW/CRM. Avoid a broad short until disclosures identify material inference-cost exposure; the key catalyst is the next two earnings cycles.
  • Maintain a measured long bias to MSFT/GOOGL/AMZN versus pure-play application AI exposure over 6-18 months, but do not assume premium API pricing persists. Reduce if cloud-management commentary indicates inference workloads are migrating materially off their platforms or AI capex monetization guidance deteriorates.

More News

From AllMind Research

Browse all research