← Back to scans

Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?

0x3fb3dfa7e6332d421dfb49d735e05b4de1a0f9b7af42d3e4b22f56dfcd1f2858 · Science and Technology · 2026-08-21
33%
Agent
60%
Market Price
-27.0%
Edge
44%
Confidence
Volume: 26,923
Spread: 10.0c
Days to resolution: 132
Markets in event: 5
Final Rationale
The named resolution source (agi.safe.ai / Scale AI) is the binding constraint: as of August 2026 it still lists claude-opus-4-7 at 36.2% no-tools, with the overall official frontier only ~46-47%, and it has consistently lagged and undercut self-reported figures by 15-20pp. Applying that observed discount to Opus 5's self-reported 56.3% implies an eventual official score in the high-30s/low-40s, not ≥60%, and the with-tools 64.7% figure is likely a red herring since the official board reports closed-book/no-tools numbers. A YES therefore requires a compound conditional — a new (or existing) Claude scoring ≥60% under the stricter official protocol AND Scale AI posting it before Dec 31 — against a backdrop of decelerating, non-monotonic benchmark progress. I weight the mechanistic trend/lag evidence (8-15%) more heavily than the thin, highly volatile Polymarket 60% anchor, but keep meaningful upside for definitional looseness (headline/third-party scores being accepted) and a possible Q4 flagship, landing at 0.33 — below both forecasters and well below the market.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 12$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-12 43% 50% 43%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news code_execution wikipedia
Sub-questions (Fermi decomposition)
  1. What is the highest HLE accuracy currently listed for any Anthropic Claude model (e.g., Claude Opus 4.5 / Claude 4.x) on the official agi.safe.ai leaderboard, and does that figure reflect no-tools or tool/search-augmented evaluation?
  2. What is the highest HLE accuracy achieved by ANY model (OpenAI, Google, xAI) as of early 2026, and how far above Claude's best is it?
  3. How fast has the top HLE score risen over time (Jan 2025 ~3-9%, mid-2025 ~20-25%, late 2025/early 2026 ~?), i.e., what is the monthly/annual rate of improvement?
  4. How frequently does Safe.ai update the agi.safe.ai leaderboard, and does it promptly add new Anthropic model releases?
  5. What Claude model releases are expected during 2026 (Claude 5 / Opus 5) and have Anthropic or credible benchmarks reported HLE numbers for them?
  6. What do sibling Polymarket markets on other HLE thresholds (e.g., 40%, 50%, 70%) and 'highest score by any model' imply about the implied distribution over Claude's 2026 peak score?
Planner reasoning
This is a Polymarket question about whether any Anthropic Claude model hits ≥60% HLE accuracy on the agi.safe.ai leaderboard during 2026. The key empirical inputs are: the current best Claude HLE score on that specific leaderboard, the overall frontier ceiling (any model), and the rate of improvement over 2024-2026, plus whether the leaderboard reports tool-augmented or no-tools numbers. Market price on Polymarket (and sibling threshold markets like 40%/50%/70%) gives an implied distribution to anchor on.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.1s 1 ## This Market's Polymarket Data **Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 60.00% - 7-day price change: -4.00% - 30-day price change: +15.00% - Total volume: $26,923 (USD notional) - Price range:
polymarket_related OK 3.1s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'Claude': 0 markets | keyword 'benchmark': 0 markets
kalshi_related OK 2.9s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Anthropic Claude': ok
claude_news OK 28.6s 10 Here are the key findings on Humanity's Last Exam (HLE) leaderboard status relevant to Anthropic Claude models in 2026: - **Official Scale AI leaderboard (partnered with Center for AI Safety, feeds agi.safe.ai)**: As of ~August 2026, top scores were gemini-3.1-pro-preview (thinking high) at 46.44%
claude_news OK 24.1s 10 Based on the research, here's the trajectory of Anthropic Claude scores on Humanity's Last Exam through 2026: - **Early 2025 baseline**: When first released in early 2025, state-of-the-art models scored only a few percent – GPT-4o managed just 2.7% and Claude 3.5 Sonnet 4.1% . - **Early 2026 (pre
gdelt_news OK 154.1s 20 GDELT: 20 articles across 3 queries (lookback=120d). "Humanity's Last Exam benchmark score": 10 hits | 'Claude Opus HLE benchmark': 10 hits | 'Anthropic Claude new model benchmark record': error GDELT rate-limited after retries (429)
code_execution OK 97.6s 0 ## Key Findings **Historical HLE trajectory (text-only, no-tool-use scores):** - Frontier (best of any lab) rose from ~9.1% (o1, Jan 2025) → ~20.3% (o3/Gemini 2.5 Pro, Apr 2025) → ~25.2–25.4% (Grok 4, GPT-5, Jul–Aug 2025) — growth is clearly **decelerating**, not exponential. - Claude-specific best
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
3. Evidence Brief Sonnet · 7649 chars
# Current state Anthropic's most recent flagship, Claude Opus 5 (released 2026-07-24), reportedly scores 56.3% on HLE without tools and 64.7% with tools (self-reported/aggregator-corroborated). However, the *official* resolution source (agi.safe.ai, which mirrors the Scale AI/CAIS leaderboard) still shows Claude's best confirmed no-tools score in the 34–46% range as of August 2026 (latest listed model: claude-opus-4-7 at 36.20%), suggesting the official leaderboard lags newer releases and/or uses a stricter/different scoring methodology than third-party trackers. Whether the market resolves YES hinges on whether the official leaderboard eventually posts a Claude no-tools score ≥60% before 2026-12-31. # Timeline of key events - 2025-01: HLE launches; SOTA models score only 3–9% (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (Wikipedia/claude_news). - 2025-02: Claude 3.7 Sonnet ~8.8% — reported. - 2025-05: Claude 4 Opus (extended thinking) ~23.8%; frontier (any model) ~20.3% — reported. - 2025-08: Frontier best (any lab) ~25.2–25.4% (Grok 4/GPT-5); Claude Opus 4.1 regresses to ~18.5% — reported (plateau signal). - 2025-12: Claude Opus 4.5 no-tools score trails Gemini 3 Pro by ~7pp (implying high-20s/low-30s%) — reported (Vellum). - 2026-02/03: Official Scale AI leaderboard: Gemini 3 Pro 37.5%, Claude Opus 4.6 Thinking Max 34.4%, GPT-5 Pro 31.6% — confirmed (IntuitionLabs). - 2026-05: Claude Opus 4.8 released; mid-to-high 30s HLE range for Anthropic's top models per trackers — reported. - 2026-06: "Claude Fable 5" reported at 55.5% (no-tools, third-party) — reported, model naming unconfirmed as official Anthropic branding at time of report. - 2026-07-24: Claude Opus 5 launched; HLE 56.3% no-tools / 64.7% with tools — reported (Anthropic/aggregators), corroborated across sources. - 2026-08 (early): Official Scale AI leaderboard still tops out with claude-opus-4-7 at 36.20% (no-tools); text-only variant shows claude-opus-4-6-thinking-max as best Claude entry — confirmed, indicating leaderboard lag vs. Opus 5 release. - 2026-08-11/19: Third-party trackers (Artificial Analysis, PricePerToken, BenchLM) list Claude Fable 5 (55.5%) and Claude Opus 5 (54.9–64.7% depending on protocol) as top HLE performers overall — reported, methodology/naming caveats apply. # Event Will any Anthropic Claude model reach ≥60% HLE accuracy (per agi.safe.ai / Scale AI official leaderboard) by 2026-12-31? # Outcomes to forecast Yes / No # Kalshi market anchor No direct Kalshi price returned for this ticker; Polymarket (same event) shows current YES price 60% (down 4% over 7 days, up 15% over 30 days), volume $26.9K, range 17%–64.5% over 29 days — treat as primary cross-market anchor given absence of Kalshi-native price in the tool output. # Sub-question answers 1. **Highest official Claude HLE score & protocol** — Official Scale AI leaderboard (Aug 2026): claude-opus-4-7 at 36.20% (no-tools); earlier Feb/Mar 2026 snapshot had Claude Opus 4.6 Thinking Max at 34.4%. All confirmed official figures are no-tools/closed-book. [claude_news/Scale AI] 2. **Best score by any model, gap to Claude** — Gemini 3.1 Pro Preview leads officially at ~46–47% (Aug 2026); GPT-5.4-pro close behind at ~44–45%. Claude trails by ~8–10pp on the official leaderboard. [claude_news] 3. **Rate of improvement** — Frontier rose 9%→20%→25% (Jan–Aug 2025), then plateaued/regressed (Opus 4.1 fell vs Opus 4); by early-mid 2026 official frontier only reached ~37–46%. Growth is decelerating, not exponential. [code_execution, claude_news] 4. **Leaderboard update cadence** — Scale AI's official leaderboard appears to lag new releases by weeks-to-months: Opus 5 (released 2026-07-24, self-reported 56.3% no-tools) is not yet reflected in the Aug 2026 official snapshot, which still shows opus-4-7 at 36.2%. [claude_news, inferred] 5. **Expected 2026 Claude releases** — Confirmed released in 2026: Opus 4.6/4.7/4.8, Claude Mythos, Claude Fable, Claude Sonnet 5, and Claude Opus 5 (2026-07-24). Self-reported/aggregator HLE for Opus 5: 56.3% no-tools, 64.7% with tools. Naming/versioning beyond Opus 5 unconfirmed for H2 2026. [claude_news, Wikipedia] 6. **Sibling markets / implied distribution** — No true HLE sibling markets found on Polymarket or Kalshi (0 matches); only loosely related Kalshi markets (Anthropic IPO odds, US equity stake market) exist, offering no direct calibration. [polymarket_related, kalshi_related] # Key facts (high-confidence, factual) 1. [claude_news/Scale AI] Official leaderboard best Claude (no-tools) ≈36–46% as of mid-2026, below 60%. 2. [claude_news] Anthropic Opus 5 (2026-07-24) self-reported/aggregator no-tools HLE = 56.3%; with-tools = 64.7%. 3. [code_execution] Logistic-fit extrapolation of official trend data implies ~20% ceiling by end-2026 under status-quo growth; naive exponential fit is invalid (unbounded). 4. [Wikipedia] HLE is a 2,500-question benchmark; no mention of Claude exceeding 60% in encyclopedic record as of last update. # Cross-market signals - Kalshi related: No direct HLE/Claude benchmark markets found; adjacent Anthropic markets (IPO odds 93%, US stake market 16%) unrelated to capability trajectory. - Polymarket: Same-event price at 60% YES, volatile (17%–64.5% range), recent 30-day rise (+15%) suggests momentum toward YES likely tied to Opus 5's 56.3%/64.7% headline figures, but 7-day pullback (-4%) suggests reassessment of official-leaderboard lag/methodology risk. - Sportsbook implied: N/A. # Analyst opinions and speculation - claude_news synthesis: resolution likely hinges on protocol (no-tools vs with-tools) and which leaderboard source is used; under with-tools, 60% already met (64.7%), but strict official no-tools leaderboard has not confirmed this. - code_execution Monte Carlo: raw P(≥60%) estimate ~8-15% under conservative ceiling assumptions purely from official-leaderboard trend extrapolation — much lower than Polymarket's 60%, reflecting divergent read of "official" vs "real-world" Claude capability. # Directional lean per outcome - **Yes**: Opus 5 already at 56.3% no-tools (self-reported) — only 3.7pp from threshold; further Anthropic releases (Opus 5.x, Sonnet 6, etc.) plausible before Dec 2026; with-tools score already exceeds 60% (64.7%) creating ambiguity that could favor YES if "HLE Accuracy" on leaderboard includes an agentic/tool-use column reported ≥60%. - **No**: Official Scale/CAIS leaderboard (the explicit resolution source) still shows sub-40% Claude scores as of Aug 2026, well below Opus 5's self-reported number — indicating either measurement discrepancy or delayed/never-added entries; historical HLE growth has been decelerating/plateauing (Opus 4.1 regression); trend-based statistical models put probability in single-to-low-double digits. # Gaps / unknowns - Unclear whether/when Scale AI will add Opus 5 (or later models) to the official leaderboard, and whether its official score will match the 56.3% self-reported figure or be lower (as historical official vs third-party gaps suggest). - No clarity on whether "HLE Accuracy" per market rules could refer to a tool-augmented score if that becomes the leaderboard's primary displayed metric. - No confirmed Claude releases/scores for Sept–Dec 2026. # Calibration anchors - Polymarket YES price (anchor): 60% (volatile, 17%-64.5% range over 29 days). - Statistical/trend-based model estimate: ~8-15% (independent of market sentiment). - Precedent: Opus 4.1 regressed vs Opus 4 on HLE, illustrating non-monotonic single-model progress; official leaderboards have consistently lagged and undercut headline/self-reported figures by ~15-20pp in 2026 snapshots.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.32
Yes 55%
No 45%
The decisive uncertainty is resolution-source strictness rather than raw capability: Claude Opus 5 (2026-07-24) is reported at 56.3% no-tools and 64.7% with tools, so a broad or tool-inclusive reading of 'highest score achieved' already clears 60%, while the official agi.safe.ai/Scale AI leaderboard still tops out near 36.2% for claude-opus-4-7 as of August 2026 due to multi-month lag and stricter methodology. Historical evidence cuts both ways — frontier HLE growth decelerated in 2025 (Opus 4.1 even regressed), and trend extrapolation on official figures yields only ~8-15%, but the 2026 official frontier already reached ~46% and Anthropic typically ships another flagship or point release in Q4, giving multiple shots at closing a ~4pp gap. The Polymarket anchor of 60% YES (up 15pp over 30 days on Opus 5 headlines, but off 4pp in the last week on leaderboard-lag reassessment) reflects exactly this split. I settle slightly above even at 0.55, discounting the market modestly because the named resolution source has consistently lagged and undercut self-reported figures by 15-20pp.
gpt-5.5
0.56
Yes 38%
No 62%
Anchoring on the cross-market Polymarket YES price of 60%, the market appears to be pricing the Opus 5 headline score heavily. The official-leaderboard evidence points toward No because the best confirmed Claude score is still only about 36.2%, and even the best official model overall is only around 46–47%, implying a large methodology or update gap. The Opus 5 self-reported 56.3% no-tools score points toward Yes because it is only 3.7pp below the threshold, and its 64.7% with-tools score creates protocol ambiguity, but the official Scale/CAIS source has not yet reflected those numbers. Decelerating benchmark progress and no confirmed post-Opus-5 release before year-end justify moving materially below the 60% market anchor, though not to the 8–15% trend-only estimate because a near-threshold model already exists.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts drift toward the market anchor (55% and 38%) without fully reconciling with the brief's own trend-based estimate of 8-15%, which is explicitly derived from the resolution source itself (official Scale AI leaderboard) — the brief flags this divergence as a "much lower" independent read, and neither forecaster explains why they weight market sentiment so much more heavily than a model built directly on the resolution criterion. 2. Neither forecast adequately engages with the "historical official vs third-party gaps" fact — that official snapshots have "consistently lagged and undercut headline/self-reported figures by ~15-20pp in 2026." Applying that discount directly to Opus 5's 56.3% self-reported score implies an eventual official score in the ~36-41% range, far below 60%, which should pull both forecasts (especially Forecast 1's 55%) much lower. 3. Both forecasts treat "another flagship in Q4" as a path to YES, but the brief notes "decelerating, not exponential" growth and a regression precedent (Opus 4.1 underperforming Opus 4) — neither forecaster quantifies how much additional gain a hypothetical Q4 release would need (they'd need to beat Opus 5 by 3.7pp+ *and* have Scale AI actually post it before Dec 31, a two-stage conditional that compounds uncertainty rather than simply "another shot at closing the gap." 4. Forecast 1's 0.55 seems to over-index on the tool-augmented 64.7% figure creating "ambiguity that could favor YES," but the brief's event definition explicitly cites agi.safe.ai/Scale AI as the resolution source, which reports only no-tools/closed-book scores per sub-question 1 — this tool-use ambiguity is likely a red herring that Forecast 1 overweights relative to Forecast 2. 5. Neither forecaster explicitly addresses the leaderboard-update-cadence risk as a distinct gating event: even if Opus 5's true capability exceeds 60%, Scale AI must both (a) add the model and (b) confirm a matching or higher score, and (b) is uncertain if official methodology is stricter (as the 34-46% vs 56% gap suggests) — this is a compound probability that should push both estimates toward the lower end of the code-execution's 8-15% band rather than the 38-55% range. 6. Both forecasts anchor heavily on the volatile Polymarket price (range 17%-64.5% over 29 days, thin $26.9K volume) without discounting for its low liquidity and demonstrated instability, which weakens its reliability as a calibration anchor relative to the more mechanistically-grounded trend/leaderboard-lag evidence.
6. Aggregator Opus · final calibrated forecast
33% (44% confidence)
Yes: 33% No: 67%
The named resolution source (agi.safe.ai / Scale AI) is the binding constraint: as of August 2026 it still lists claude-opus-4-7 at 36.2% no-tools, with the overall official frontier only ~46-47%, and it has consistently lagged and undercut self-reported figures by 15-20pp. Applying that observed discount to Opus 5's self-reported 56.3% implies an eventual official score in the high-30s/low-40s, not ≥60%, and the with-tools 64.7% figure is likely a red herring since the official board reports closed-book/no-tools numbers. A YES therefore requires a compound conditional — a new (or existing) Claude scoring ≥60% under the stricter official protocol AND Scale AI posting it before Dec 31 — against a backdrop of decelerating, non-monotonic benchmark progress. I weight the mechanistic trend/lag evidence (8-15%) more heavily than the thin, highly volatile Polymarket 60% anchor, but keep meaningful upside for definitional looseness (headline/third-party scores being accepted) and a possible Q4 flagship, landing at 0.33 — below both forecasters and well below the market.
Pipeline Timing
Total pipeline time: 266.5s
Per-tool research timings shown in the Research section above.