← Back to scans

Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?

0xcb251bef81c1eaa9267a4bc8b9811922722d9afc9a6e7f5477b2167884011ba2 · Science and Technology · 2026-08-30
60%
Agent
74%
Market Price
-13.5%
Edge
56%
Confidence
Volume: 23,866
Spread: 3.0c
Days to resolution: 123
Markets in event: 6
Final Rationale
The official leaderboard sits at 46.44%, requiring a steep +13.5-point gain in ~4 months on a benchmark explicitly designed to resist saturation, with the growth trajectory visibly decelerating. The critique correctly flags that both forecasts under-weighted the verification-lag mechanism (unofficial 55-65% scores are entirely self-reported, 0 verified) and the direct base-rate analog of analysts' failed '50% by end-2025' prediction, plus the thin, momentum-driven Polymarket proxy price deserves heavier discounting. However, frontier capability does appear near the threshold per multiple third-party trackers, and a single verified Q4 frontier release could clear 60% given historical 10-15pt jumps per release cycle. I pull below both forecasts toward the logistic-fit midpoint, landing at 60% YES — meaningfully below the 73.5% market anchor but respecting the genuine upside from continued rapid iteration.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 3$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-22 65% 72% 54%
2026-08-15 40% 56% 45%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' shown on the official leaderboard at agi.safe.ai, and which model holds it?
  2. What has been the trajectory of top HLE scores over time (e.g., from ~3-9% in Jan 2025 to the latest state-of-the-art), and what does extrapolating that trend imply for end of 2026?
  3. Have any recently announced models (e.g., GPT-5.x, Gemini 3, Grok 4/5, Claude) reported HLE scores near or above 60%, including with tool-use/agentic setups, and would those count under the leaderboard's 'HLE Accuracy' metric?
  4. How frequently and promptly does the official agi.safe.ai leaderboard update with new frontier model results, and does it include tool-augmented configurations?
  5. What frontier model releases are expected in 2026 that could plausibly push HLE scores above 60%?
  6. What probabilities do related prediction markets assign to adjacent HLE thresholds (e.g., ≥50%, ≥70% in 2026) or other AI benchmark milestone markets, and are those consistent with this market's price?
Planner reasoning
This is a Polymarket question on whether any model reaches ≥60% on the official Humanity's Last Exam leaderboard by end of 2026. The key drivers are the current top score, the historical rate of improvement on HLE, upcoming frontier model releases, and whether the official leaderboard (agi.safe.ai) is actively updated with new results. I'll anchor on the Polymarket price, triangulate with related markets (including sibling threshold markets like 50%/70%), and research recent HLE results and leaderboard behavior.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 73.50% - 7-day price change: +0.00% - 30-day price change: +12.50% - Total volume: $23,866 (USD notional) - Price range: 45.00% - 73.50% - Data point
polymarket_related OK 1.1s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 1 markets | keyword 'AI benchmark 2026': 0 markets
kalshi_related OK 1.0s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok
claude_news OK 27.1s 12 Here are the key findings from research (as of late August 2026): - **Official Scale AI/CAIS leaderboard (agi.safe.ai / scale.com):** Per a recent Wikipedia snapshot, the top scores listed are Gemini 3.1 Pro Preview at 46.44% accuracy, followed by GPT-5.4 Pro at 44.32%, Muse Spark at 40.56%, and C
gdelt_news OK 116.6s 30 GDELT: 30 articles across 4 queries (lookback=60d). "Humanity's Last Exam score": 10 hits | 'HLE benchmark AI model': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Read timed out. (read timeout=30) | "Gemini 3 Humanity's Last Exam": 10 hits | "GPT-5 Humanity's Last Exam": 10 hits
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 58.0s 0 **Key quantitative findings** - **Data used (score, month index from Jan‑2025):** Jan‑25 Deep Research ≈9%, Jun‑25 Grok‑4 ≈25‑44%, Aug‑25 GPT‑5 ≈25‑42%, Nov‑25 Gemini‑3 ≈37‑45%. Target date Dec‑2026 = month 24. - **Linear OLS extrapolation (all 4 points):** predicted Dec‑2026 score = **67% (low/no
3. Evidence Brief Sonnet · 6709 chars
# Current state The official resolution source (agi.safe.ai / Scale AI leaderboard) shows a top HLE Accuracy of **~46.44%** (Gemini 3.1 Pro Preview, as of 2026-08-16), still well below the 60% threshold. Third-party trackers (Artificial Analysis, llm-stats.com) report unofficial/self-reported scores of 55–65% for newer Claude "Opus 5"/"Fable 5" variants, but these are NOT reflected on the official leaderboard and are explicitly flagged as unverified self-reported results. # Timeline of key events - 2025-01: HLE launched; SOTA models scored single digits (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (Wikipedia). - 2025-06: Grok-4 (with tool use) reported ~38–45% — reported (mixed sourcing on tools vs. no-tools). - 2025-08: GPT-5 reported ~25–42% depending on configuration — reported. - 2025-11: Gemini 3.0 Pro reached ~37–45%, official "highest recorded" cited as ~38.3% — reported. - 2026 (early): Official leaderboard scores clustered ~38–40% — reported. - 2026-07-09: Meta's "Muse Spark 1.1" benchmark claims disputed amid conflict-of-interest allegations — reported/rumored. - 2026-07-10: GPT-5.6 Sol/Terra/Luna launched — confirmed release, HLE scores not specified in research. - 2026-08-16: Official leaderboard snapshot: Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Claude Opus 4.6 (Thinking) 34.44% — confirmed (Wikipedia/IntuitionLabs). - 2026-08 (ongoing): Artificial Analysis reports Claude Fable 5 (55.5%) and Claude Opus 5 (54.9%) — reported, unofficial tracker, not on agi.safe.ai. - 2026-08 (ongoing): llm-stats.com reports Claude Opus 5 at 64.7%, but flags "0 verified / 101 self-reported" results — rumored/unverified. # Event Will the highest score on the official Humanity's Last Exam leaderboard (agi.safe.ai) reach ≥60% accuracy by Dec 31, 2026? # Outcomes to forecast - Yes (≥60% achieved by any model on official leaderboard) - No (highest score stays below 60%) # Kalshi market anchor No direct Kalshi YES price was returned by kalshi_direct; only a cross-listed Polymarket price for the identical event is available: **73.5% YES** (as of latest snapshot), up from 45% a month ago (+12.5% 30-day trend, flat 7-day). Volume is thin ($23,866 total over 39 days), suggesting a fairly illiquid, still-forming consensus. This should be treated as the best available proxy anchor given the missing native Kalshi quote. # Sub-question answers 1. **Current highest official HLE score** — Gemini 3.1 Pro Preview at 46.44% leads the official Scale AI/agi.safe.ai leaderboard as of 2026-08-16 [claude_news/Wikipedia/IntuitionLabs]. 2. **Trajectory of top scores** — Roughly 3–9% (Jan 2025) → ~25–45% (mid/late 2025) → ~38–46% (2026 to date). Growth has been rapid but decelerating; no model has yet crossed 50% on the *official* board [claude_news, code_execution]. 3. **Recent models near/above 60%** — Not on the official leaderboard. Third-party sites (Artificial Analysis: Claude Fable 5 ~55.5%, Opus 5 ~54.9%; llm-stats.com: Opus 5 ~64.7%) report higher figures, but these are self-reported/unverified and use possibly different methodology (tool-use, adaptive reasoning) not confirmed to match "HLE Accuracy" as officially defined [claude_news]. 4. **Leaderboard update cadence** — Updates continuously as new frontier models are verified, but with lag and a private held-out set to detect overfitting; it does not automatically absorb third-party/self-reported tool-augmented results [claude_news]. 5. **Expected 2026 releases** — GPT-5.6 (Sol/Terra/Luna), Gemini 3.7 Flash, DeepSeek V4-Pro already released or imminent per GDELT; none yet confirmed to push official HLE above 50–55% [gdelt_news]. 6. **Adjacent market pricing** — No native Kalshi ≥50%/≥70% threshold markets found; Polymarket price for this exact 60% threshold is 73.5%, implying markets lean Yes but with recent upward momentum, not settled certainty. # Key facts (high-confidence, factual) 1. [Wikipedia/claude_news] Official leaderboard top score ≈46.44% (Gemini 3.1 Pro Preview) as of 2026-08-16. 2. [claude_news] HLE launched Jan 2025 with SOTA scores in single digits. 3. [claude_news] llm-stats.com explicitly notes 0 of 101 evaluated results are "verified" — all self-reported. 4. [polymarket_direct] Polymarket price for this exact question: 73.5% YES, up sharply (+12.5pts) over 30 days. 5. [code_execution] Trend-extrapolation models disagree widely: linear fits imply 67–95%; logistic (saturating) fits imply 50–68%, straddling the 60% line. # Cross-market signals - Kalshi related: No directly comparable AI-benchmark markets found; unrelated hits (TV show release, SCOTUS case) not informative. - Polymarket: Same-event price 73.5% YES, with strong 30-day upward momentum (45%→73.5%), thin volume (~$24k) — moderate confidence signal. - Sportsbook implied: N/A. # Analyst opinions and speculation - Mid-2025 analysts predicted "50% by end of 2025" — did not materialize on official board [claude_news]. - Some commentary suggests official/third-party score gap (contamination-control private set, verification lag) may cause official scores to lag true frontier capability by many months, which could delay a Yes resolution even if capability effectively exceeds 60%. - Code-execution quantitative synthesis leans mildly Yes (~55-60% probability) but flags extreme sensitivity to model choice/tools-inclusion. # Directional lean per outcome - **Yes**: Rapid historical growth curve (9%→46% in ~20 months); multiar third-party trackers already claim 55–65%; frontier labs (GPT-5.6, Gemini 3.7, Claude Opus 5) continuing rapid iteration; Polymarket pricing 73.5%. - **No**: Official leaderboard still 14+ points below threshold with only ~4 months of 2026 lapsed at last snapshot; unofficial high scores are unverified/self-reported and may not count; logistic/saturation models suggest diminishing returns as HLE was designed to resist saturation; benchmark stewards periodically raise difficulty via private held-out sets. # Gaps / unknowns - No direct Kalshi native quote retrieved — only Polymarket cross-price used as proxy anchor. - Unclear whether official leaderboard will incorporate/verify the high third-party scores (Claude Opus 5, 54-65%) before Dec 2026. - No confirmed frontier-lab roadmap data for Q4 2026 releases. # Calibration anchors - Cross-market anchor (Polymarket, same event): 73.5% YES, rising. - Official leaderboard current top score: 46.44%, needs +13.5 points in ~4 months to resolve Yes — a steep but not unprecedented jump given historical ~10-15pt gains per 3-6 months. - Base rate precedent: HLE scores rose ~35-40 points in 20 months; achieving another 13+ points in 4 months is plausible but requires continued rapid pace.
4. Forecasters Opus + GPT · in parallel
claude-fable-5
0.55
Yes 68%
No 32%
The official agi.safe.ai leaderboard sits at 46.44% (Gemini 3.1 Pro Preview, Aug 2026), needing +13.5 points in ~4 months — steep but consistent with the historical 10-15pt gains per 3-6 month cycle. Third-party trackers already report Claude Opus 5/Fable 5 at 55-65%, suggesting frontier capability is at or near the threshold, though these are unverified and may use non-comparable methodology. The Polymarket proxy anchor is 73.5% YES with strong upward momentum; I discount modestly below that because the official leaderboard's verification lag and private held-out set could delay resolution even if capability crosses 60%, and quantitative logistic fits straddle the 60% line. Analysts' prior 'rapid crossing' predictions (50% by end-2025) also underdelivered on the official board.
gpt-5.6-sol
0.58
Yes 64%
No 36%
The 73.5% Polymarket YES price is the best available consensus anchor, although thin volume and the missing direct Kalshi quote reduce its reliability. The official leader is only at 46.44%, so reaching 60% requires a substantial 13.5-point gain in roughly four months amid signs of slowing benchmark progress. Unofficial Claude results around 55–65% indicate that the necessary capability may already be near, but their self-reported, methodologically uncertain status creates meaningful verification and leaderboard-lag risk. Balancing rapid historical gains and prospective frontier-model updates against that official-score gap supports a moderately lower YES probability than the market proxy.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters anchor heavily on the Polymarket 73.5% price despite explicit brief warnings that it's thin-volume ($23,866 over 39 days), only a proxy (not native Kalshi), and had a sharp +12.5pt run-up in 30 days — this could reflect speculative momentum/manipulation rather than settled information, yet neither forecast meaningfully discounts for that possibility beyond a token few points. 2. Neither forecast grapples with the base-rate failure precedent explicitly flagged in the brief: "Mid-2025 analysts predicted '50% by end of 2025' — did not materialize on official board," which is a direct analog to the current situation (analysts/trackers claiming near-threshold scores that haven't yet been verified) and should push probability down more than ~10-15 points from the naive trend extrapolation. 3. Both treat the "13.5 points in 4 months" gap symmetrically with historical "10-15pt gains per 3-6 months," but understate that HLE's design explicitly resists saturation via private held-out sets and periodic difficulty increases — the brief notes recent trajectory is "decelerating" (38-46% growth has slowed vs. earlier 25-45% jumps), which argues for a logistic/saturating curve, not linear continuation. 4. Neither forecaster addresses the verification-lag mechanism concretely: even if a model's true capability exceeds 60%, the official leaderboard requires confirmed/verified submission, and with llm-stats.com showing "0 of 101 verified," there's no established base rate for how long verification takes — this operational bottleneck could mechanically cap official 2026 resolution regardless of underlying capability, and deserves more explicit weight than a single mitigating sentence. 5. Both forecasts converge tightly (0.64-0.68) despite the brief's own quantitative synthesis showing wide disagreement ("linear fits imply 67-95%; logistic fits imply 50-68%, straddling the 60% line") — this model uncertainty warrants a wider, more uncertain probability (perhaps closer to 50-55%) rather than both independently landing comfortably above 60% YES. 6. Neither forecast considers the tail scenario where a benchmark provider disputes/discredits a high score before verification (as already happened with "Muse Spark 1.1... disputed amid conflict-of-interest allegations" in July 2026), suggesting institutional friction/skepticism toward inflated scores could further delay or block a Yes resolution even if raw capability is there.
6. Aggregator Opus · final calibrated forecast
60% (56% confidence)
Yes: 60% No: 40%
The official leaderboard sits at 46.44%, requiring a steep +13.5-point gain in ~4 months on a benchmark explicitly designed to resist saturation, with the growth trajectory visibly decelerating. The critique correctly flags that both forecasts under-weighted the verification-lag mechanism (unofficial 55-65% scores are entirely self-reported, 0 verified) and the direct base-rate analog of analysts' failed '50% by end-2025' prediction, plus the thin, momentum-driven Polymarket proxy price deserves heavier discounting. However, frontier capability does appear near the threshold per multiple third-party trackers, and a single verified Q4 frontier release could clear 60% given historical 10-15pt jumps per release cycle. I pull below both forecasts toward the logistic-fit midpoint, landing at 60% YES — meaningfully below the 73.5% market anchor but respecting the genuine upside from continued rapid iteration.
Pipeline Timing
Total pipeline time: 209.3s
Per-tool research timings shown in the Research section above.