
Survey of 157 enterprise respondents finds an “evaluation gap”: 50% have shipped an AI/LLM agent or feature that passed internal evaluations but then failed a customer-facing workflow, including 25% that have seen it happen more than once. Trust is low—only 5% fully trust automated evaluations, with the largest complaint (29%) that evals don’t align with real-world outcomes—and production monitoring is mostly limited to uptime/cost (51% monitor only functioning vs 23% monitoring whether outputs are correct). Despite this, 66% already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are engineering toward it within 12 months (33%), while 64% plan to adopt or switch evaluation platforms within a year.
This is less a “better testing” story than a budget reallocation story: enterprises are monetizing AI before they can fully verify it, which shifts spend toward runtime monitoring, human escalation, and audit trails rather than pure eval tooling. That favors platforms with deep distribution into logs/metrics/security and workflow controls — think DDOG, SNOW, CRWD, and the cloud providers’ native stacks — while standalone eval vendors risk being commoditized by integration friction and price pressure.
The second-order loser is any customer-facing software vendor that is using agents to suppress headcount or support costs without building robust review loops. In the next 1-3 quarters, the risk is not “AI doesn’t work”; it’s margin leakage from exception handling, rework, and incident response that shows up as higher opex before revenue upside arrives. A high-profile failure in regulated verticals would accelerate procurement for governance/observability tools and slow zero-human deployment in finance, healthcare, and retail, even if the broader autonomy trend remains intact.
Contrarian view: the market may be underestimating how sticky the trust gap is for point solutions, but overestimating how much it slows adoption overall. If enterprises keep shipping despite weak confidence, the winners are not necessarily the “best evals” vendors — they are the vendors that become the control plane for oversight. For GAP specifically, this has no direct fundamental read-through; treat it as zero-alpha unless the company discloses meaningful customer-facing AI automation exposure.
AI-powered research, real-time alerts, and portfolio analytics for institutional investors.
Overall Sentiment
mildly negative
Sentiment Score
-0.25
Ticker Sentiment