AI models flub these intelligence tests. Can you fare any better?
Source: MIT Technology Review
The article highlights rapid AI progress on puzzle benchmarks—e.g., models improving from solving ~18% of NYT Connections puzzles (late 2024) to near-perfect accuracy (early 2025). It also emphasizes persistent capability gaps, especially in spatial/3D manipulation and visual reasoning (e.g., weaknesses seen in benchmarks like ARC-AGI), along with issues like over-reliance on memorized patterns when puzzles resemble training data. Overall, the piece is informational about AI strengths and failure modes rather than a catalyst for immediate financial market repricing.
Analysis
The investable takeaway is not that models are “solved”; it is that benchmark wins are increasingly cheap while real monetization still depends on workflow integration, multimodality, and error containment. That shifts capital toward platform vendors that can combine text, vision, retrieval, and device-native inference rather than pure model branding. In that frame, GOOGL is better positioned than most because it can commercialize capability gains across search, cloud, and assistant surfaces; AAPL has longer-dated optionality if spatial/visual understanding becomes a real consumer feature, but this is a product-cycle story, not a near-term earnings driver.
For content and puzzle franchises, the second-order risk is gradual substitution rather than immediate disruption: if users can offload casual cognitive games to assistants, engagement may erode over 6-18 months even if brand awareness rises. That matters for NYT only if games materially support retention and bundle pricing; otherwise the impact is mostly sentiment noise. The contrarian view is that the market tends to overrate benchmark progress and underrate edge-case brittleness: failures in visual/spatial reasoning suggest enterprise buyers will still pay for guardrails, which supports software vendors with human-in-the-loop products more than standalone AI claims.
Near term, this reads as a low-conviction tape signal. The most actionable move is relative, not directional: own the AI platform with the broadest distribution and most monetization paths, and fade consulting-heavy AI narratives that rely on generic benchmark improvement. Falsifiers are simple: if GOOGL’s AI product adoption fails to show up in cloud/search monetization over the next 1-2 quarters, or if NYT game engagement holds despite broader assistant usage, then the thesis weakens materially.
AllMind Terminal
AI-powered research, real-time alerts, and portfolio analytics for institutional investors.
Request TrialMarket Sentiment
Overall Sentiment
neutral
Sentiment Score
0.10
Ticker Sentiment
Key Decisions for Investors
- Long GOOGL / short IBM as a 3-6 month relative-value expression on whose AI stack can monetize multimodal reasoning; exit if GOOGL cloud/search AI revenue does not inflect by the next earnings cycle or IBM bookings reaccelerate.
- No immediate trade in AAPL; keep as a watch item for WWDC and device-level AI features tied to vision/spatial tasks. Reassess only if management demonstrates a consumer feature that changes upgrade cadence over the next 1-2 product cycles.
- Monitor NYT for game/engagement deceleration over the next 2 quarters rather than shorting outright; if Games DAUs or bundle retention soften while AI usage rises, consider a small hedge against slower subscriber growth.
- Avoid forcing a directional trade in GAP/TSCC/TSTS on this note; the article has no clear revenue or margin linkage for retail/specialty retail names, so the risk/reward is too weak to justify capital.
- Set an alert on GOOGL/IBM earnings and AI product commentary; if the market starts rewarding benchmark headlines without corresponding revenue conversion, fade the move with call spread sellers rather than outright shorts.
More News
- Wall Street is pitching data centers as a major real estate bet. The risks are piling up
- Global PC shipments crater 20% as rising prices hammer demand
- Cramer: Investors selling Apple on latest iPhone 18 report are 'stupid as plywood'
- Sequoia-backed Nuvacore seeks $2.5bn valuation in funding round
- Chick-fil-A is keeping humans at the drive-thru while McDonald’s tests AI ordering and defends a separate pricing algorithm in court
- Apple reportedly cuts iPhone 18 Pro component orders by at least 15%
From AllMind Research
- Anthropic IPO Preview: Valuation, Timing, and What to Watch
- Shein After the IPO: Venue, Valuation, and What Must Be Proved
- What AI Research Tools Should a Small Hedge Fund Buy First?
- AI Research Tools for Pension Funds and Allocators
- AllMind's Data Standardization Methodology: Our Approach to Fundamentals