← Back to scans

Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?

0xcb251bef81c1eaa9267a4bc8b9811922722d9afc9a6e7f5477b2167884011ba2 · Science and Technology · 2026-08-22
65%
Agent
72%
Market Price
-7.5%
Edge
54%
Confidence
Volume: 23,163
Spread: 3.0c
Days to resolution: 131
Markets in event: 6
Final Rationale
The resolution-relevant official CAIS/Scale leaderboard sits at ~46.4% as of Aug 2026, requiring ~14 points in four months — above the trailing pace and well above the early-2026 deceleration of ~1.3 pts/mo — which is the strongest case for No. However, the official site is demonstrably lagging rather than capping: independent no-tools trackers already report 55–56% (Fable 5, Opus 5), so a single official refresh plus one more frontier launch (GPT-6/Gemini 4-class/Deep Think update) could clear 60% without any re-acceleration in true capability. The 72.5% Polymarket anchor deserves discounting because it is only a proxy for the actual ticker, may resolve on looser 'any reported score' wording, and moved on tool-assisted BenchLM headlines whose methodology (and self-reported-benchmark CoI concerns) is disputed. Weighing the critique's points on resolution ambiguity, institutional update cadence, and the Monte Carlo's wider/lower 45–55% band, I settle modestly below both the market and Forecast 1 but above the conservative scenario: 65% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 11$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-15 40% 56% 45%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' listed on the official agi.safe.ai leaderboard, and by which model?
  2. How has the SOTA HLE accuracy progressed month-by-month from January 2025 to now (e.g., o3 ~20%, GPT-5 ~25-42%, Gemini 3 Pro ~37.5%, Deep Think ~41%+), and what is the implied monthly/annual improvement rate?
  3. Does the official leaderboard record tool/search-augmented scores (which are substantially higher, e.g. 45%+) or only text-only no-tools scores, and would a tool-augmented run count for resolution?
  4. Which frontier model releases are expected before Dec 31, 2026 (GPT-5.5/6, Gemini 3.5/4, Claude next, Grok 5, DeepSeek/Qwen), and have any labs pre-announced HLE results near or above 60%?
  5. Is the agi.safe.ai leaderboard still being actively updated in 2026, or has it stalled/been superseded (which would affect resolution)?
  6. What do adjacent prediction markets (Polymarket/Kalshi AI benchmark markets, other HLE thresholds like 50%/70%) imply about the probability distribution of the 2026 HLE max score?
Planner reasoning
This is a Polymarket AI-benchmark question, so the market price is the primary anchor and the key empirical inputs are the current HLE leaderboard SOTA, the rate of improvement over the past 12-24 months, and whether the official agi.safe.ai leaderboard counts tool-augmented runs. I'll pull the direct market, search news for the latest HLE records and upcoming frontier model releases, and use code execution to extrapolate the trend line to Dec 2026.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.1s 1 ## This Market's Polymarket Data **Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 72.50% - 7-day price change: +11.50% - 30-day price change: +27.50% - Total volume: $23,163 (USD notional) - Price range: 45.00% - 72.50% - Data poin
polymarket_related OK 3.5s 2 Scanned 100 active Polymarket markets, kept 2 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 2 markets | keyword 'AI benchmark': 0 markets | keyword 'best AI model': 0 markets
kalshi_related OK 3.3s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'artificial intelligence model': no matches
claude_news OK 29.3s 12 Based on research findings: - **Original benchmark difficulty (early 2025 baseline):** When the benchmark was first released in early 2025, leading AI models scored in the single digits: GPT-4o managed just 2.7%, Claude 3.5 Sonnet reached 4.1%, and OpenAI's o1 model achieved only 8.0% . Source: ht
claude_news OK 23.4s 8 Here are key findings on Humanity's Last Exam (HLE) performance in 2026: - **Official Scale AI/CAIS leaderboard (labs.scale.com) as of ~Aug 2026** still shows comparatively lower scores, topped by gemini-3.1-pro-preview (thinking high) at 46.44%, followed by gpt-5.4-pro-2026-03-05 at 44.32% — wel
gdelt_news OK 117.5s 20 GDELT: 20 articles across 3 queries (lookback=120d). "Humanity's Last Exam benchmark score": 10 hits | 'HLE accuracy record AI model': error GDELT rate-limited after retries (429) | 'agi.safe.ai leaderboard': 10 hits
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 130.8s 0 ## Key Findings — HLE Max Score ≥60% by Dec 31, 2026 **Historical trend (running SOTA, monthly):** - Rapid phase (Jan–Aug 2025): 9% → 13% → 20% → 26–42%, i.e., ~2.7–3.3 pts/month average slope (deterministic linear fits: no-tools slope ≈2.73%/mo, with-tools slope ≈3.27%/mo) - Recent deceleration (A
3. Evidence Brief Sonnet · 6715 chars
# Current state As of Aug 2026, the **official CAIS/Scale AI leaderboard (agi.safe.ai / labs.scale.com)** — the market's stated resolution source — tops out at **~46.4%** (Gemini 3.1 Pro preview), still below the 60% threshold and still being actively updated. Independent trackers (Artificial Analysis, BenchLM, PricePerToken) using different protocols (adaptive reasoning, "max effort," or tool-assisted runs) report Claude-family models scoring 53–65%, but it is unconfirmed whether these match the official site's no-tools "HLE Accuracy" metric used for resolution. # Timeline of key events - 2025-01: HLE released; baseline scores near-zero (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0%) — confirmed (IntuitionLabs/Wikipedia). - 2025-mid: o3-based Deep Research ~26.6% — reported. - 2025-11-18: Gemini 3 Pro 37.5% (no tools); Gemini 3 Deep Think 41.0% (no tools) — confirmed (Google blog). - 2025-11 (pre-Gemini 3): GPT-5 Pro top score 31.6% — reported (AOL/press). - 2026-02-12: Gemini 3 Deep Think updated to 48.4% (no tools), official Google claim — confirmed (Google blog). - 2026-06: Claude Fable 5 launches, 53% HLE — reported (Artificial Analysis). - 2026-07-09: Meta's Muse Spark 1.1 HLE claims disputed as conflict-of-interest — reported (Digg). - 2026-07-24: Claude Opus 5 — 56.3% no-tools / 64.7% with tools — reported (MarkTechPost). - 2026-08 (~mid): Official Scale/CAIS leaderboard shows Gemini 3.1 Pro (46.44%) and GPT-5.4 Pro (44.32%) as top two — reported (claude_news synthesis). - 2026-08-11: PricePerToken leaderboard: Fable 5 55.5%, Opus 5 54.9%, GPT-5.6 Sol 49.5% — reported. - 2026-08-22: BenchLM (third-party, tool-inclusive) shows Opus 5 64.7%, Mythos 5 64.5%, Muse Spark 1.1 62.1% — reported, methodology disputed. # Event Will the highest score on any model on the official Humanity's Last Exam leaderboard (agi.safe.ai) reach ≥60% "HLE Accuracy" by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No direct kalshi_direct quote was returned for this ticker; the ticker string matches a Polymarket condition ID and the only direct price data available is from **polymarket_direct**: current YES/"≥60%" price **72.5%**, up sharply from 45% a month ago (+27.5% in 30d, +11.5% in 7d), volume ~$23k. Treat this as the closest available consensus anchor pending confirmed Kalshi data — it has moved strongly toward Yes recently, likely driven by Claude Opus 5/BenchLM headlines. # Sub-question answers 1. **Current highest official HLE score** — Official Scale/CAIS leaderboard tops at ~46.44% (Gemini 3.1 Pro preview, thinking-high), with GPT-5.4 Pro at 44.32%, as of Aug 2026 (claude_news). 2. **Month-by-month SOTA progression** — Roughly 9%→41% (Jan–Nov 2025, ~2.7–3.3 pts/mo), then a marked slowdown to ~1.3 pts/mo through early 2026 (48.4% Feb 2026), before third-party trackers jump to 53–65% by mid-2026 via Claude Fable 5/Opus 5 (code_execution Monte Carlo synthesis). 3. **Tools vs. no-tools on leaderboard** — Official leaderboard historically reports no-tools scores; the 60%+ figures (Opus 5 64.7%) are explicitly tool-augmented per MarkTechPost/BenchLM, while no-tools scores remain mid-50s (Artificial Analysis 55.5%). Ambiguous whether official site would ever list tool-augmented numbers as "HLE Accuracy." 4. **Expected frontier releases before Dec 2026** — GPT-5.5/5.6, Claude Fable 5/Opus 5/Mythos 5, Gemini 3.1 Pro/Deep Think updates, Meta Muse Spark all already released by mid-2026 per timeline; further GPT-6/Gemini 4/Grok 5 releases plausible before year-end (code_execution notes 2-3 more major releases expected). 5. **Is agi.safe.ai still active?** — Yes, confirmed still updated with new entries as of Aug 2026 (claude_news), though it lags behind third-party trackers. 6. **Adjacent markets** — No other Kalshi HLE-threshold markets found (kalshi_related returned irrelevant matches). Polymarket price for this same question is 72.5% and rising fast. # Key facts (high-confidence, factual) 1. [Google blog] Gemini 3 Deep Think reached 48.4% (no tools) on HLE by Feb 2026 — official lab claim. 2. [claude_news/Scale] Official CAIS/Scale leaderboard top score ~46.44% (Gemini 3.1 Pro) as of Aug 2026 — below 60%. 3. [MarkTechPost] Claude Opus 5: 56.3% no-tools, 64.7% with tools (Jul 2026). 4. [Artificial Analysis] Independent no-tools tracker's current max is 55.5% (Claude Fable 5, Aug 2026) — still under 60%. 5. [Wikipedia] HLE co-created by CAIS and Scale AI; agi.safe.ai is CAIS's official site. # Cross-market signals - Kalshi related: no directly comparable HLE-threshold markets found. - Polymarket: same-question price 72.5%, up from 45% a month ago — strong recent Yes momentum. - Sportsbook implied: n/a. # Analyst opinions and speculation - code_execution Monte Carlo: blended P(≥60%) ≈ 45–55%, but highly scenario-dependent (18% conservative vs 97% optimistic) depending on whether Aug–Nov 2025 slowdown reflects a real ceiling near 55–65% or a temporary lull. - claude_news: resolution "likely hinges on which leaderboard/methodology is treated as authoritative" — official site lags well behind tool-assisted third-party claims. - Digg commentary flags conflict-of-interest concerns about self-reported benchmark scores (Meta Muse Spark), raising doubt about inflated third-party numbers generally. # Directional lean per outcome - **Yes**: Rapid 2025 trajectory, multiple 50%+ no-tools claims by mid-2026 (Fable 5, Opus 5), momentum in adjacent Polymarket price (72.5%), several more frontier releases expected before Dec 2026. - **No**: The specific resolution source (official agi.safe.ai/Scale leaderboard) sits well below 60% (~46%) as of Aug 2026 and has shown decelerating no-tools progress; ambiguity whether tool-assisted 60%+ scores would even count; strict resolution rules favor the stricter, lagging official metric. # Gaps / unknowns - No confirmed live snapshot of the exact agi.safe.ai number at time of forecast (only ~Aug 2026 secondhand reports). - Unclear whether official leaderboard will ever adopt tool-augmented scoring before Dec 2026 close. - True Kalshi YES price for this ticker not directly retrieved (only Polymarket data available, though ticker matches). - Four months (Sep–Dec 2026) of further progress unaccounted for; could close gap from 46%→60% if trend continues at even 2-3 pts/mo. # Calibration anchors - Polymarket price for same question: 72.5% (rising fast). - Official/resolution-relevant score trajectory: ~9% (Jan 2025) → ~46% (Aug 2026), ~2 pts/mo average over 19 months; needs another ~14 pts in ~4 months to hit 60% via official metric — faster than trailing-12-month average pace, but consistent with historical jump-driven accelerations at major model launches.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.52
Yes 73%
No 27%
The official CAIS/Scale leaderboard — the stated resolution source — sits at ~46.4% (Gemini 3.1 Pro) as of Aug 2026, requiring ~14 points in four months, faster than the trailing ~2 pts/mo average. However, third-party no-tools trackers already report 55.5–56.3% (Claude Fable 5, Opus 5) and tool-augmented runs exceed 60% (Opus 5 64.7%, BenchLM 62–65%), so the official leaderboard appears to lag rather than reflect a true ceiling, and several more frontier launches (GPT-6, Gemini 4-class, Deep Think updates) are expected before year-end. Ambiguity over whether tool-assisted scores count and whether the lagging official site updates in time is the main downside risk, plus conflict-of-interest doubts about self-reported numbers. The Polymarket anchor on the identical question is 72.5% and rising sharply (+27.5 pts in 30d), which I largely accept given the reported no-tools scores are already within striking distance.
gpt-5.5
0.56
Yes 58%
No 42%
The closest market anchor is Polymarket at 72.5% Yes and rising, but I discount it because the resolution-relevant official CAIS/Scale leaderboard is only around 46.4% as of Aug 2026, well below the 60% threshold. The current official score points toward No because reaching 60% requires roughly a 14-point gain in four months on the stricter no-tools-like metric, while third-party 60%+ results appear tool-assisted or methodologically disputed. The SOTA progression and expected frontier releases point toward Yes because multiple independent no-tools claims are already in the mid-50s and a single major model jump could plausibly clear 60% before year-end. Overall, I stay below the 72.5% consensus anchor due to official-leaderboard lag and methodology risk, but still favor Yes given rapid recent frontier progress and several remaining release opportunities.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor heavily on the 72.5% Polymarket price without adequately weighting that it's not confirmed to be the same resolution source as Kalshi (the brief explicitly notes "no direct kalshi_direct quote was returned" and the ticker "matches a Polymarket condition ID" — this is a proxy, not a confirmed match), so treating it as a strong consensus anchor may be miscalibrated. 2. Neither forecast grapples seriously with the "decelerating trend" evidence: official no-tools progress slowed from ~2.7-3.3 pts/mo (2025) to ~1.3 pts/mo (early 2026), yet both assume a re-acceleration to ~14 pts in 4 months is plausible without justifying why the deceleration would reverse rather than continue. 3. The code_execution Monte Carlo synthesis explicitly gave a much wider and lower blended estimate (45-55%, with an 18% conservative scenario) — both forecasters land above this range (58% and 73%) without explaining why they reject the more conservative scenario weighting. 4. Neither forecast adequately addresses resolution ambiguity risk as a distinct probability mass: if the official site's methodology never adopts tool-assisted scoring and remains lagging by design, "No" could resolve almost mechanically regardless of frontier capability — this deserves more explicit downweighting of Yes than either gives, especially forecast 1's 73%. 5. Forecast 1's rationale doesn't sufficiently discount the CoI concerns raised about Meta Muse Spark self-reported scores, which the brief flags as casting "doubt about inflated third-party numbers generally" — this should reduce confidence in the 55-65% third-party figures broadly, including Opus 5 and Fable 5 claims, more than either forecast reflects. 6. Both forecasts converge on similar reasoning (official lag vs. third-party progress) but diverge substantially in output (58% vs 73%) without a clear articulation of what specific evidence drives the 15-point gap — this inconsistency itself suggests the true uncertainty band is wider than either point estimate implies, and a probability nearer the lower/conservative end (or with wider variance) may be warranted given four months is a short window for a leaderboard that updates on its own institutional cadence.
6. Aggregator Opus · final calibrated forecast
65% (54% confidence)
Yes: 65% No: 35%
The resolution-relevant official CAIS/Scale leaderboard sits at ~46.4% as of Aug 2026, requiring ~14 points in four months — above the trailing pace and well above the early-2026 deceleration of ~1.3 pts/mo — which is the strongest case for No. However, the official site is demonstrably lagging rather than capping: independent no-tools trackers already report 55–56% (Fable 5, Opus 5), so a single official refresh plus one more frontier launch (GPT-6/Gemini 4-class/Deep Think update) could clear 60% without any re-acceleration in true capability. The 72.5% Polymarket anchor deserves discounting because it is only a proxy for the actual ticker, may resolve on looser 'any reported score' wording, and moved on tool-assisted BenchLM headlines whose methodology (and self-reported-benchmark CoI concerns) is disputed. Weighing the critique's points on resolution ambiguity, institutional update cadence, and the Monte Carlo's wider/lower 45–55% band, I settle modestly below both the market and Forecast 1 but above the conservative scenario: 65% Yes.
Pipeline Timing
Total pipeline time: 232.9s
Per-tool research timings shown in the Research section above.