← Back to scans

Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?

0x3fb3dfa7e6332d421dfb49d735e05b4de1a0f9b7af42d3e4b22f56dfcd1f2858 · Science and Technology · 2026-08-30
33%
Agent
58%
Market Price
-25.5%
Edge
46%
Confidence
Volume: 33,405
Spread: 13.0c
Days to resolution: 123
Markets in event: 5
Final Rationale
The named resolution source (agi.safe.ai/Scale AI) currently shows the best Claude entry at ~36–40%, roughly 20+ points short of the 60% threshold, while the frontier-wide official no-tools SOTA is only ~47% and has advanced ~10 points in 9 months — making a Claude leap to 60%+ on the official board within 2026 a genuine outlier scenario. Anthropic's self-reported Opus 5 numbers (56.3% no-tools / 64.7% with tools) already clear the bar, but independent replication (~53%) and the documented vendor-to-leaderboard discount pattern (36.2% listed vs. 64.7% claimed) suggest listing Opus 5 would likely still land short of 60%. The main YES pathways are therefore (a) resolution being adjudicated on self-reported/aggregator figures rather than the official leaderboard, or (b) a further generation (Opus 6-class) landing and being listed above 60% before year-end — each plausible but not dominant. The thin, volatile Polymarket price (58.5%, $33K volume, 17–65% range) most likely reflects conflation of self-reported with-tools claims with the actual resolution source, so I discount it substantially toward the leaderboard-anchored Monte Carlo range (25–36%), landing at ~0.33 Yes — below both prior forecasts, per the critique's stronger case.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 4$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 33% 60% 44%
2026-08-12 43% 50% 43%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' listed for any Anthropic Claude model on the official agi.safe.ai leaderboard, and what did Anthropic itself report for Claude Opus 4.5 / Sonnet 4.5 on Humanity's Last Exam?
  2. Does the agi.safe.ai leaderboard report no-tools scores or tool-augmented (search + code) scores, and how large is the gap between those two regimes for frontier models?
  3. What is the overall SOTA HLE score as of early 2026 (e.g., Gemini 3 Deep Think, GPT-5.x), and how fast did the frontier improve over the prior 12 months?
  4. How quickly and how consistently has the agi.safe.ai leaderboard been updated with new model releases (i.e., is there a risk a strong Claude model exists but is never listed)?
  5. What Anthropic model releases are expected in 2026 (Claude 5 / Opus 5 cadence), and what HLE gains have successive Claude generations delivered?
  6. What do sibling Polymarket markets on other HLE thresholds (40%, 50%, 70%, or other labs) imply about the crowd's distribution over 2026 Claude HLE scores?
Planner reasoning
This is a Polymarket AI-benchmark question, so the market price on Polymarket is the primary anchor, plus sibling markets in the same event series (other thresholds like 40%/50%/70%) which reveal the implied distribution. The key empirical inputs are: what the agi.safe.ai leaderboard currently reports for Claude models, the historical rate of HLE score improvement through 2025 (roughly 3% → ~40%+ for frontier models with tools), whether the leaderboard scores are tool-augmented or no-tools, and Anthropic's expected 2026 model release cadence.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will the highest score achieved by an Anthropic Claude model on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 58.50% - 7-day price change: -5.50% - 30-day price change: +7.00% - Total volume: $33,405 (USD notional) - Price range:
polymarket_related OK 2.9s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'Claude': 0 markets | keyword 'Anthropic': 0 markets | keyword 'AI benchmark 2026': 0 markets
kalshi_related OK 2.7s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Anthropic Claude': ok
claude_news OK 27.7s 13 Here are the key findings on HLE performance for Claude models as of 2026: **Official/independent leaderboard (Scale AI, hosted via agi.safe.ai) — much lower scores:** - The official Scale AI HLE leaderboard (linked from agi.safe.ai) shows, as of mid-late August 2026, the text-only no-tools leaderb
claude_news OK 17.5s 8 Based on my research, here are the key findings: - **Question already resolved YES**: As of July 24, 2026, Anthropic's Claude Opus 5 (part of the "Claude 5" generation) scored 56.3% without tools and 64.7% with them on Humanity's Last Exam , well above the 60% threshold when tools are used. Source
gdelt_news OK 117.3s 20 GDELT: 20 articles across 3 queries (lookback=90d). "Humanity's Last Exam benchmark score": 10 hits | "Claude Opus Humanity's Last Exam": 10 hits | 'Anthropic new model benchmark 2026': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?q
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 65.4s 0 ## Key Findings **Model setup:** Logistic (log-odds) progress curve, current Claude best ≈ 10–20% HLE, frontier SOTA ≈ 40–50% (tool-augmented), implied frontier growth ≈ 3.0 logits/year (from ~3%→40-50% in 2025), Monte Carlo over 4 uncertain factors (200k-400k draws). - **Base-case probability P(C
3. Evidence Brief Sonnet · 7328 chars
# Current state The market resolves on the **official agi.safe.ai / Scale AI HLE leaderboard**, which as of August 2026 shows Anthropic's best Claude model (claude-opus-4-6-thinking-max / Opus 4.6, or Opus 4.7 in combined mode) topping out around **36–40%** — well below 60%. Anthropic's own self-reported model-card and third-party aggregator numbers (Artificial Analysis, llm-stats, BenchLM) claim much higher "with tools" scores for Claude Opus 5 (56.3% no-tools / 64.7% with tools, per Anthropic's July 2026 release), but these are NOT yet reflected on the resolution-source leaderboard, creating a critical protocol/source discrepancy. # Timeline of key events - 2025 (throughout): HLE scores near-zero for early models (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0%) — confirmed, historical baseline [intuitionlabs.ai]. - 2025-11: Gemini 3 scores 37.5% no-tools / 45.8% with tools on HLE — confirmed via multiple aggregators [sentisight.ai]. - 2026-06-10/12: Claude Fable 5 released; reported as "smartest" model but noted media caution about benchmark comparability — confirmed release, reported framing [livemint.com, timesofindia.com]. - 2026-07-24/25: Claude Opus 5 released; Anthropic self-reports 56.3% (no tools) / 64.7% (with tools) on HLE — reported, per Anthropic materials via MarkTechPost, corroborated by llm-stats.com and BenchLM.ai (confirmed as *self-reported/aggregator* figures, not yet on official Scale AI leaderboard). - 2026-07-09: Commentators flag conflict-of-interest concerns re: Meta Muse Spark's HLE claims — reported, signals broader benchmark-gaming skepticism [digg.com]. - 2026-08 (mid-late): Official Scale AI/agi.safe.ai leaderboard (text-only, no-tools) shows top Claude entry (Opus 4.6-thinking-max) below Gemini 3.1 Pro (47.31%) and GPT-5.4-Pro (45.32%); combined (with-tools) leaderboard shows top Claude (Opus 4.7) at 36.20% — confirmed, primary resolution source. - Ongoing: Independent verification (Artificial Analysis) of Opus 5 without Anthropic's own protocol puts scores closer to 53% — reported, suggests self-reported 64.7% may not replicate under standardized/independent conditions. # Event Will any Anthropic Claude model reach ≥60% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct data was returned; the only direct market pricing available is **Polymarket: 58.5% YES** (down 5.5% over 7 days, up 7.0% over 30 days, range 17–65%, $33.4K volume, 38 data points). Treat this as the crowd anchor in absence of Kalshi data; note it is notably higher than what official-leaderboard evidence supports. # Sub-question answers 1. **Current highest official Claude HLE score vs. Anthropic's own reporting** — Official Scale AI/agi.safe.ai leaderboard: Claude tops out ~36–40% (Opus 4.6/4.7). Anthropic's Opus 4.6 system card claims "leads all frontier models" at 40.0% no-tools; Opus 5 (July 2026) self-reports 56.3% no-tools/64.7% with tools [claude_news, Anthropic materials]. 2. **No-tools vs. tool-augmented leaderboard structure** — agi.safe.ai/Scale AI hosts both a text-only no-tools leaderboard and a combined (with-tools) leaderboard; the gap for frontier models is large (~10-20 points), e.g., Gemini 3: 37.5%→45.8% with tools [claude_news]. 3. **Overall SOTA as of early-to-mid 2026** — Gemini 3.1 Pro leads no-tools official leaderboard at 47.31%, GPT-5.4-Pro at 45.32% (Aug 2026); frontier moved from ~37.5% (Nov 2025) to ~47% (Aug 2026) — a slower pace than the 2025 near-zero-to-40% jump [claude_news]. 4. **Leaderboard update cadence/risk of Claude being unlisted** — Leaderboard appears actively maintained (multiple Claude and competitor entries current through Aug 2026), so omission risk is low, but there's a lag: newer models (e.g., Opus 5) may not yet appear on official rankings despite public release. 5. **2026 Anthropic release cadence** — Opus 4.6 (mid-2026) → Opus 5 (Jul 2026) → Fable 5/Mythos (mid-2026, restricted). Successive generations reportedly deliver large jumps in self-reported (non-official) HLE scores (30.8%→40.0%→56.3% no-tools across Opus 4.5→4.6→5). 6. **Sibling markets/crowd distribution** — No Polymarket sibling markets found (0 matches for HLE/Claude/Anthropic keywords). Kalshi-related markets are unrelated (IPO race, ERISA, TV release). No cross-market distributional signal available. # Key facts (high-confidence, factual) 1. [claude_news] Official Scale AI leaderboard: top Claude no-tools ~36-40%, combined/tools ~36.2% (Opus 4.7), as of Aug 2026 — far below 60%. 2. [claude_news/MarkTechPost] Anthropic self-reports Opus 5 at 56.3% (no tools)/64.7% (with tools) — not yet confirmed on official leaderboard. 3. [Artificial Analysis via trendingtopics.eu] Independent re-evaluation of Opus 5 shows ~53%, below Anthropic's self-reported figure. 4. [polymarket_direct] Market price 58.5%, volatile (range 17-65%), suggesting substantial disagreement/uncertainty. 5. [code_execution model] Monte Carlo estimate: P(YES) ≈ 25-36%, highly sensitive to whether resolution counts tool-augmented scores and whether Claude tracks frontier pace. # Cross-market signals - Kalshi related: no directly relevant markets found (IPO race, ERISA, unrelated). - Polymarket: only this market itself; no sibling threshold markets (40%/50%/70%) found to triangulate distribution. - Sportsbook implied: none applicable. # Analyst opinions and speculation - claude_news synthesis flags this as fundamentally a **protocol/definition dispute**: if resolution uses Anthropic's own "with tools" self-reported numbers, threshold already met; if using official Scale AI no-tools/combined leaderboard (the stated resolution source), threshold is far from met (~36-40%). - code_execution model’s calibrated point estimate (~25-30%) explicitly discounts self-reported Anthropic figures given Claude's documented historical underperformance vs. OpenAI/Google on HLE specifically. # Directional lean per outcome - **Yes**: Anthropic's own Opus 5 claims (64.7% with tools) already exceed threshold; rapid 2025-2026 generational jumps (30.8%→56.3% no-tools in ~1 year) suggest further gains plausible; Polymarket prices it fairly likely (58.5%). - **No**: The actual official resolution source (agi.safe.ai/Scale AI) shows Claude far below 60% (~36-40%) even in tool-augmented mode; independent verification of Opus 5 undercuts Anthropic's self-reported 64.7%; frontier-wide official no-tools leaderboard SOTA is only ~47% (non-Claude); benchmark saturation/deceleration risk. # Gaps / unknowns - Whether Scale AI/agi.safe.ai will list Opus 5 or later models with tool-augmented configurations matching Anthropic's internal setup, and whether resolved scores will replicate the 64.7% claim. - No Kalshi-direct price was returned in research; anchor relies on Polymarket only. - No visibility into planned Claude 5-series (Opus 6?) roadmap beyond Opus 5 for remainder of 2026. # Calibration anchors - Polymarket YES price (used as anchor in absence of Kalshi data): 58.5%. - Model-based estimate: ~25-36% (code_execution Monte Carlo), reflecting weight on official/no-tools leaderboard evidence. - Precedent: benchmarks (e.g., MMLU, GPQA) often show 6-12 month lag between vendor self-reported SOTA claims and independent/official leaderboard confirmation.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.31
Yes 45%
No 55%
This hinges almost entirely on resolution protocol: Anthropic's self-reported Opus 5 figures (56.3% no-tools / 64.7% with tools, July 2026) already clear 60% in the tool-augmented configuration, while the stated resolution source — the official agi.safe.ai/Scale AI leaderboard — shows the best Claude entry at only ~36–40% (Opus 4.6/4.7), far below threshold, with frontier no-tools SOTA (Gemini 3.1 Pro) at just 47.31%. Independent re-evaluation by Artificial Analysis puts Opus 5 nearer 53%, undercutting the 64.7% claim and matching the historical 6–12 month lag between vendor claims and leaderboard confirmation. The Polymarket anchor at 58.5% (volatile, 17–65% range, thin $33K volume) appears to price the self-reported/with-tools reading, whereas the code-execution Monte Carlo anchored on official leaderboard data gives 25–36%. I split the difference, leaning modestly to No because the explicitly named resolution source currently shows Claude ~20+ points short even in combined mode, though further 2026 releases and a possible leaderboard update of Opus 5-class models keep Yes live.
gpt-5.6-sol
0.61
Yes 36%
No 64%
The 58.5% Polymarket YES price is the crowd anchor, but the official resolution leaderboard currently places Claude around 36–40%, substantially below 60%. Although Anthropic reports Opus 5 at 64.7% with tools, its no-tools result is 56.3% and independent testing is closer to 53%, pointing toward No under standardized evaluation. The actively maintained leaderboard may add Opus 5 or another 2026 release, and Anthropic's rapid release cadence leaves a meaningful path to Yes, especially if a comparable tool-augmented run is accepted. However, official frontier performance is only about 47% and has recently advanced more slowly, so I discount the market anchor in favor of No.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor heavily on Polymarket (58.5%) despite the brief explicitly noting this is a thin, volatile market (17-65% range, $33K volume) likely mispriced due to conflating self-reported "with tools" figures with the actual resolution source—yet neither forecast heavily discounts this anchor's reliability, and Forecast 1 essentially "splits the difference" toward it (0.45) rather than deferring to the stronger leaderboard evidence. 2. Neither forecast adequately weighs the "gap-closing pace" sub-question: no-tools official SOTA moved from ~37.5% (Nov 2025) to ~47% (Aug 2026), a ~10-point gain in 9 months across the whole frontier, meaning Claude would need to jump from ~40% to 60%+ (20 points) in the remaining 4-5 months of 2026 while also leapfrogging Gemini/GPT-5.4 — both treat the trajectory as plausible without stress-testing this specific quantitative extrapolation. 3. Both forecasts underweight the leaderboard's demonstrated *lag/discount* pattern: even if Opus 5 is added, the with-tools independent estimate (~53%, per Artificial Analysis) is below 60%, and the combined leaderboard already shows Anthropic's actual accepted score is far lower than self-reported (36.2% vs 64.7% claimed) — implying resolution likely requires a further generation beyond Opus 5, which neither forecast explicitly models as a separate scenario (e.g., "Opus 6 releases and clears 60%" vs "Opus 5 gets listed and clears 60%"). 4. Neither critique addresses the specific ambiguity of which leaderboard variant (no-tools vs. combined) actually governs resolution — the brief flags this as "a fundamental protocol/definition dispute," yet both forecasters implicitly assume the combined/with-tools leaderboard is the relevant comparison point without confirming this is the stated official criterion, which could materially shift the answer if no-tools is definitive (currently only ~40% for Claude, further from threshold). 5. The model-based Monte Carlo anchor (25-36%) is cited by both, but Forecast 2's final 36% and Forecast 1's 45% show inconsistent treatment of this anchor — Forecast 1 barely moves toward it while Forecast 2 lands almost exactly within its range, yet neither explains why the crowd anchor should pull the estimate up as far as it does given the brief's clear steer that official-leaderboard evidence dominates. 6. Neither forecast considers the possibility that agi.safe.ai/Scale AI could revise methodology or add a separate "tool-augmented" track that retroactively validates something close to Anthropic's self-reported number before year-end — a tail scenario that could resolve YES via a protocol change rather than a raw performance leap, which is a distinct pathway from "Claude actually improves."
6. Aggregator Opus · final calibrated forecast
33% (46% confidence)
Yes: 33% No: 67%
The named resolution source (agi.safe.ai/Scale AI) currently shows the best Claude entry at ~36–40%, roughly 20+ points short of the 60% threshold, while the frontier-wide official no-tools SOTA is only ~47% and has advanced ~10 points in 9 months — making a Claude leap to 60%+ on the official board within 2026 a genuine outlier scenario. Anthropic's self-reported Opus 5 numbers (56.3% no-tools / 64.7% with tools) already clear the bar, but independent replication (~53%) and the documented vendor-to-leaderboard discount pattern (36.2% listed vs. 64.7% claimed) suggest listing Opus 5 would likely still land short of 60%. The main YES pathways are therefore (a) resolution being adjudicated on self-reported/aggregator figures rather than the official leaderboard, or (b) a further generation (Opus 6-class) landing and being listed above 60% before year-end — each plausible but not dominant. The thin, volatile Polymarket price (58.5%, $33K volume, 17–65% range) most likely reflects conflation of self-reported with-tools claims with the actual resolution source, so I discount it substantially toward the leaderboard-anchored Monte Carlo range (25–36%), landing at ~0.33 Yes — below both prior forecasts, per the critique's stronger case.
Pipeline Timing
Total pipeline time: 222.6s
Per-tool research timings shown in the Research section above.