← Back to scans

Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?

0x3fb3dfa7e6332d421dfb49d735e05b4de1a0f9b7af42d3e4b22f56dfcd1f2858 · Science and Technology · 2026-08-12
43%
Agent
50%
Market Price
-7.5%
Edge
43%
Confidence
Volume: 19,337
Spread: 1.0c
Days to resolution: 141
Markets in event: 5
Final Rationale
Both forecasters essentially anchored to Polymarket's 50.5% and nudged up, without independently pricing the conjunction required for Yes: the official agi.safe.ai leaderboard must (a) actually update with a Claude entry before Dec 31 and (b) display an accuracy of ≥60%. The strongest documented facts push the other way: the official snapshot still shows only Opus 4.7 at ~36.2% with a confirmed multi-month lag, the primary leaderboard column is no-tools where Anthropic's own best self-report is 56.3% (short of 60%), and third-party recomputes have historically come in below Anthropic's self-serving framing (Opus 4.6: 40.0% recompute vs 'leads all frontier models' claim), with grader revisions an additional downside risk. Paths to Yes remain real — a with-tools column at 64.7%, a post-Opus-5 release (Fable/Mythos variants), and a genuinely fast-moving benchmark — which is why I don't go far below the market. Scenario weighting (roughly 55-65% chance of a timely official Claude update × ~55-65% chance that update reads ≥60%, plus a modest with-tools/wildcard boost) lands just under the coin-flip anchor.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 22$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct claude_news claude_news gdelt_news polymarket_related kalshi_related wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' figure for any Anthropic Claude model on the official Humanity's Last Exam leaderboard (agi.safe.ai), and as of what date/model (e.g., Claude Opus 4.5)?
  2. What is the current overall SOTA HLE accuracy across all labs (e.g., Gemini 3 Pro/Deep Think, GPT-5.x), and does the official leaderboard list tool/search-augmented scores or text-only no-tools scores?
  3. What has been the month-over-month rate of increase in SOTA HLE accuracy since the benchmark's January 2025 launch, and what does linear/exponential extrapolation imply for Anthropic's best score by Dec 31, 2026?
  4. How frequently does the agi.safe.ai leaderboard get updated with new Anthropic models, and is there a lag or omission risk for Claude releases?
  5. What is Anthropic's expected 2026 model release cadence (Claude 5 / Opus 5 / Sonnet 5), and have they signaled HLE performance targets or published HLE numbers in model cards?
  6. Are there parallel prediction markets (Polymarket/Kalshi) on other HLE thresholds (e.g., 50%, 70%, or 'any model above X') whose prices imply a consistent distribution?
Planner reasoning
This is a Polymarket question about a specific benchmark threshold, so the market price is the primary anchor and the key empirical facts are (a) the current highest Claude HLE score on the agi.safe.ai leaderboard, (b) the overall SOTA trajectory of HLE scores, and (c) Anthropic's expected 2026 model release cadence. The main uncertainty is the pace of HLE improvement (and whether tool-augmented/search-enabled scores count on the leaderboard), so I want news search plus a math step to extrapolate the improvement curve.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 50.50% - 7-day price change: -2.00% - 30-day price change: +5.50% - Total volume: $19,337 (USD notional) - Price range:
claude_news OK 43.2s 9 Here are key findings on Humanity's Last Exam (HLE) leaderboard standings relevant to this forecast: **Official Scale AI / agi.safe.ai leaderboard (independent, no-tools text evaluation):** - As of ~Aug 2026, the top scores are gemini-3.1-pro-preview (thinking high) at 46.44%, gpt-5.4-pro-2026-03-
claude_news OK 34.9s 11 Based on my research, here are the key findings on Humanity's Last Exam (HLE) score trajectory and Claude model performance in 2026: - **Early 2025 baseline**: When the benchmark was first released in early 2025, leading AI models scored in the single digits: GPT-4o managed just 2.7%, Claude 3.5 S
gdelt_news OK 103.8s 30 GDELT: 30 articles across 3 queries (lookback=90d). "Humanity's Last Exam benchmark score": 10 hits | 'Claude Opus HLE accuracy benchmark': 10 hits | 'Anthropic Claude new model benchmark 2026': 10 hits
polymarket_related OK 4.3s 2 Scanned 100 active Polymarket markets, kept 2 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'Claude score': 0 markets | keyword 'benchmark': 2 markets
kalshi_related OK 4.2s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Anthropic': ok
wikipedia OK 0.1s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 57.6s 0 ## Key Findings **Trend fits to SOTA HLE milestones (Jan 2025→Nov 2025, t in months):** - Linear fit: slope ≈ **2.85 pts/month** → projects SOTA ≈ **73%** by Dec 2026 (likely too aggressive; extrapolates the recent Gemini-3-Deep-Think jump linearly) - Logistic fit (more realistic, saturating): ceil
3. Evidence Brief Sonnet · 7024 chars
# Current state The market resolves on the OFFICIAL agi.safe.ai (Scale AI/CAIS) Humanity's Last Exam leaderboard. That leaderboard, as of the most recent snapshot (~Aug 2026), shows the best-ranked Claude model (Opus 4.7) at only ~36% (no-tools) — well below 60%. However, Anthropic's own July 2026 Opus 5 launch reports a no-tools HLE score of 56.3% and a with-tools score of 64.7%, figures not yet reflected on the official leaderboard snapshot cited. This creates a live reconciliation gap between self-reported Anthropic numbers and the lagging official leaderboard that will determine resolution. # Timeline of key events - 2025-01: HLE launches; frontier models score single digits (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (intuitionlabs.ai). - 2025-11-18: Gemini 3 Pro sets record 37.4% (no-tools), surpassing GPT-5 Pro's 31.64% — confirmed (TechCrunch). - 2025-11 (Opus 4.5 launch): Claude trails Gemini 3 Pro by ~7pts no-tools, ~2pts with search — confirmed (Vellum). - 2026-02: Opus 4.6 system card claims Anthropic "leads all frontier models" on HLE; third-party recompute: 40.0% no-tools (vs GPT-5.2's 50.0%), 53.1% with tools — reported, self-serving framing flagged. - 2026-02-12: Gemini 3 Deep Think sets new SOTA, 48.4% no-tools — confirmed (MarkTechPost). - 2026-05-28/29: Claude Opus 4.8 launched (dynamic workflows) — confirmed (moneycontrol, mashable); benchmark specifics not detailed. - 2026-07-24: Claude Opus 5 launches; Anthropic reports 56.3% no-tools / 64.7% with tools on HLE — reported by Anthropic + independently corroborated (aireleasetracker.com, benchlm.ai). - 2026-08 (aggregator snapshots): Artificial Analysis/pricepertoken list "Claude Fable 5" at 55.5% and Opus 5 at 54.9%; felloai.com cites Fable 5 at 53.3% — conflicting, unverified model-naming caveat (Fable/Mythos not on Anthropic's standard Opus/Sonnet/Haiku scheme per Wikipedia; Wikipedia does confirm a 2026 "Mythos"/"Fable" release exists). - 2026-08 (official Scale/agi.safe.ai snapshot): top Claude entry still listed as Opus 4.7 at ~36.2%, i.e., official leaderboard appears NOT yet updated with Opus 5 — reported, key lag/omission risk. # Event Will any Anthropic Claude model reach ≥60% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No Kalshi-direct price returned for this specific ticker; nearest cross-market anchor is Polymarket at 50.5% YES (down 2pts/7d, up 5.5pts/30d), range 17–60.5% over 20 days, ~$19.3K volume — treat as the consensus to beat. # Sub-question answers 1. **Current highest Claude HLE figure on official leaderboard** — ~36.2% (Claude Opus 4.7, no-tools) per Scale/agi.safe.ai Aug 2026 snapshot (claude_news). Self-reported Opus 5 (Jul 2026) claims 56.3% no-tools but not yet visible on the official leaderboard. 2. **SOTA across labs / tools vs no-tools** — Official leaderboard SOTA no-tools ~46–48% (Gemini 3.1 Pro Preview, GPT-5.4 Pro); Gemini 3 Deep Think claimed 48.4%. Leaderboard is primarily no-tools; "with tools" scores exist but are typically self-reported in model cards, not the primary leaderboard column. 3. **Rate of increase / extrapolation** — Naive Fermi/code_execution fit (using data only through Nov 2025) implies Claude ≈19.7–57.2% by Dec 2026 (mean ~34%, P(≥60%)≈10-25%). This is now stale: actual Jul 2026 Opus 5 self-reported score (56.3%) already exceeds that mean, showing real trajectory outpaced the earlier linear/logistic fits. 4. **Leaderboard update lag** — Evidence of a real lag: Opus 5 (Jul 2026) not reflected in Aug 2026 official snapshot showing only Opus 4.7. Omission risk is confirmed and material for year-end resolution timing. 5. **2026 release cadence / targets** — Confirmed cadence: Opus 4.6 (Feb), Opus 4.8 (May), Sonnet 5 (mid-2026), Opus 5 (Jul 24). Wikipedia confirms additional restricted "Mythos" and public "Fable" variants released in 2026. No explicit numeric HLE target stated by Anthropic beyond claiming benchmark leadership. 6. **Parallel markets** — No other Kalshi/Polymarket HLE threshold markets found (0 keyword matches for "HLE"/"Humanity's Last Exam"); only this Polymarket contract exists as a direct comparator. # Key facts (high-confidence, factual) 1. [claude_news/Scale AI] Official leaderboard Aug 2026: top Claude entry (Opus 4.7) ≈36.2%, below Gemini/GPT SOTA (~44-46%). 2. [claude_news/Anthropic] Opus 5 (Jul 24, 2026) self-reported: 56.3% no-tools, 64.7% with tools. 3. [TechCrunch/MarkTechPost] Non-Anthropic SOTA progressed 37.4%→48.4% no-tools between Nov 2025–Feb 2026. 4. [Wikipedia] Anthropic released additional 2026 model tiers "Mythos" (restricted) and "Fable" (public), beyond standard Opus/Sonnet/Haiku naming — corroborates aggregator "Fable 5" mentions. 5. [Polymarket] Current YES 50.5%, range 17–60.5%, trending up over 30d. # Cross-market signals - Kalshi related: no direct HLE market; tangential Anthropic markets (IPO race 92%, sector classification 85%) show no benchmark-relevant signal. - Polymarket: 50.5% YES, coin-flip pricing, recent modest uptrend (+5.5% 30d) consistent with Opus 5's July release news. - Sportsbook implied: N/A. # Analyst opinions and speculation - claude_news synthesis: "if with-tools counts, threshold arguably already cleared; if only no-tools counts, Claude is just under 60% but climbing rapidly." - Digg/commentators flag conflict-of-interest concerns about self-reported benchmark claims (Meta Muse Spark case) — generalizable caution for Anthropic's own Opus 5 figures pending official leaderboard confirmation. - code_execution Fermi model (using pre-Opus5 data) is stale/underestimates given Opus 5's actual reported jump to 56.3%. # Directional lean per outcome - **Yes**: Opus 5 no-tools (56.3%) is very close to threshold; with-tools (64.7%) already exceeds it; steep recent trajectory (40%→56.3% in ~5 months) plus ~5 more months and likely further Opus/Sonnet iterations before Dec 2026 close. - **No**: Official agi.safe.ai leaderboard still shows only ~36%, with confirmed update lag/omission of Opus 5; resolution explicitly keys on "HLE Accuracy" label on official source, likely the no-tools column, and leaderboard methodology updates (grader revisions) have sometimes lowered reported scores; unverified aggregator model names ("Fable 5") add noise/risk of overstatement. # Gaps / unknowns - Whether official leaderboard will list a "with tools" HLE Accuracy figure that could trigger Yes at 64.7%. - Timing/likelihood of agi.safe.ai actually updating with Opus 5 (or later models) before Dec 31, 2026 cutoff. - True status/naming of "Claude Fable 5"/"Mythos" — possibly same models as Opus 5 variants, unclear if independently verified on official leaderboard. # Calibration anchors - Polymarket YES 50.5% (primary cross-market anchor). - Precedent: HLE SOTA has moved from single digits (Jan 2025) to ~48% (Feb 2026) to Anthropic's self-reported 56%+ (Jul 2026) — a fast-moving benchmark where large multi-month jumps are historically plausible.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.34
Yes 53%
No 47%
The Polymarket anchor sits at a near coin-flip 50.5% YES with a modest 30-day uptrend, and the underlying evidence is genuinely balanced. Pro-Yes: Anthropic's July 2026 Opus 5 self-reports 56.3% no-tools and 64.7% with tools, only a few points from (or already past) the 60% bar, with a steep trajectory (40%→56.3% in ~5 months) and further Opus/Sonnet/'Fable' iterations expected before year-end. Pro-No: the official agi.safe.ai leaderboard — the stated resolution source — still lists only Opus 4.7 at ~36.2% as of the Aug 2026 snapshot, a confirmed and material update-lag/omission risk, and the primary column is no-tools where 56.3% falls short; grader/methodology revisions have historically trimmed scores, and aggregator names like 'Fable 5' (53–55%) are unverified. I nudge marginally above the anchor because the pace of both model releases and Anthropic's own reported gains outpaces the earlier extrapolations, but the leaderboard-update bottleneck keeps this close to 50/50.
gpt-5.5
0.52
Yes 54%
No 46%
The consensus anchor is Polymarket at 50.5% YES, and I move modestly above it because Anthropic’s July 2026 Opus 5 self-reported HLE score of 56.3% no-tools is already close to 60%, while its with-tools score of 64.7% would clear the threshold if accepted. The official agi.safe.ai leaderboard still showed only Opus 4.7 at about 36.2% in the Aug 2026 snapshot, which points toward No because resolution depends on that official source and there is confirmed update/omission risk. The broader trajectory points toward Yes because Claude reportedly moved from around 40% in February to 56.3% in July, leaving several months for another model or leaderboard update before year-end. However, the likely primary leaderboard column is no-tools, and unverified aggregator scores around 53–56% suggest the 60% threshold is not yet clearly documented.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters converge near 53-54% with minimal spread, essentially anchoring to Polymarket without independently weighting the strongest structural fact: the official leaderboard has a *confirmed* multi-month update lag (Opus 5 released Jul 24, still absent from Aug snapshot), meaning the resolution source itself may simply never catch up by Dec 31 regardless of true model capability — this argues for a lower Yes than a coin-flip, not a nudge above it. 2. Neither forecast explicitly models the "with-tools" ambiguity as a distinct scenario with its own probability — if agi.safe.ai adds or already has a with-tools column showing 64.7%, that's a near-certain Yes trigger, but if it's no-tools only, Claude sits at 56.3% needing further gains; this bimodal resolution-criteria uncertainty should have been decomposed rather than blended into a single point estimate. 3. Both treat the trajectory (40%→56.3% in ~5 months) as supporting Yes, but this rate is Anthropic's *self-reported* number, not the official leaderboard's independent measurement — the leaderboard's own no-tools SOTA only moved 37.4%→48.4% (Nov 2025–Feb 2026) among all labs, a much slower pace that better predicts what agi.safe.ai will actually display by year-end. 4. Neither forecaster addresses the risk flagged in the brief that "grader/methodology revisions have historically trimmed scores" — official leaderboard re-grading is a documented downside risk that could keep even an updated Opus 5 entry below 60% even if Anthropic's self-report says 56-65%. 5. Both forecasts land within 1 point of each other (53% vs 54%), suggesting insufficient independent reasoning; given the genuine bimodality here (leaderboard-lag-driven No vs. self-reported-score-driven Yes), a wider dispersion or explicit scenario-weighted calculation (e.g., P(leaderboard updates by Dec 31) × P(score ≥60% if updated)) would have been more diagnostic than a simple anchor-and-nudge approach. 6. Neither forecast weighs the "Fable 5"/"Mythos" naming ambiguity as a potential wildcard — if these unverified aggregator entries (53-55%) are actually distinct, more advanced Claude variants that could reach 60%+ before Dec 31 and get officially listed under a different name, this represents an under-explored path to Yes that isn't captured in either rationale.
6. Aggregator Opus · final calibrated forecast
43% (43% confidence)
Yes: 43% No: 57%
Both forecasters essentially anchored to Polymarket's 50.5% and nudged up, without independently pricing the conjunction required for Yes: the official agi.safe.ai leaderboard must (a) actually update with a Claude entry before Dec 31 and (b) display an accuracy of ≥60%. The strongest documented facts push the other way: the official snapshot still shows only Opus 4.7 at ~36.2% with a confirmed multi-month lag, the primary leaderboard column is no-tools where Anthropic's own best self-report is 56.3% (short of 60%), and third-party recomputes have historically come in below Anthropic's self-serving framing (Opus 4.6: 40.0% recompute vs 'leads all frontier models' claim), with grader revisions an additional downside risk. Paths to Yes remain real — a with-tools column at 64.7%, a post-Opus-5 release (Fable/Mythos variants), and a genuinely fast-moving benchmark — which is why I don't go far below the market. Scenario weighting (roughly 55-65% chance of a timely official Claude update × ~55-65% chance that update reads ≥60%, plus a modest with-tools/wildcard boost) lands just under the coin-flip anchor.
7. Entry Decision Opus Trader
Side
NO
Position Size
$700
Contracts
1400
Trader Confidence
43%
Trade Rationale
The edge is -7.5pp (bot 43% vs market 50.5%) and technically straddles 50%, but with the market essentially at a coin flip this is only a nominal direction disagreement — the real substance is the resolution-mechanics argument. That argument is concrete and specific: the official agi.safe.ai leaderboard still shows only Opus 4.7 at ~36.2% with a documented multi-month update lag (Opus 5 released Jul 24, absent from the Aug snapshot), the primary no-tools column has Anthropic's own best self-report at 56.3% (below the 60% bar), and third-party recomputes have historically landed under Anthropic's framing. That is a real conjunction (leaderboard updates in time × updated score reads ≥60%) that a 50/50 market plausibly hasn't priced. Offsetting this: both underlying ensemble members actually came in at 53-54% (above market), so the 43% is a V2 override rather than a consensus view, and forecaster confidence is a middling 0.43 — that argues for a below-baseline position rather than a skip.
Allocation Logic
$700 — below the $1000 baseline because the bearish view is an override of an ensemble that leaned the other way (1pp spread at 53-54%), and because the book already holds a correlated Anthropic-capability NO position; the edge magnitude and the specific, verifiable leaderboard-lag evidence justify entering rather than passing.
Entry price: $0.50
Current: $0.32
Status: OPEN
P&L: -$245.00
Pipeline Timing
Total pipeline time: 234.9s
Per-tool research timings shown in the Research section above.