← Back to scans

Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?

0xcb251bef81c1eaa9267a4bc8b9811922722d9afc9a6e7f5477b2167884011ba2 · Science and Technology · 2026-08-15
40%
Agent
56%
Market Price
-16.0%
Edge
45%
Confidence
Volume: 18,131
Spread: 2.0c
Days to resolution: 138
Markets in event: 6
Final Rationale
Resolution hinges solely on the official agi.safe.ai/Scale CAIS no-tools leaderboard, which sits at ~37.5–46.4% — a required jump of ~14–23 points in roughly the final four months of 2026, larger in absolute terms than any prior comparable window and against a visibly decelerating official curve. The 55–65% figures from BenchLM/Artificial Analysis are the main Yes pathway but are unverified, plausibly tool-augmented or ensemble-based, and in one case (Muse Spark 1.1) actively disputed for conflict of interest; a contested posting adds two-sided tail risk rather than clean upside. Label-noise (~30% flawed chem/bio items) imposes a genuine ceiling, though an 'HLE-Verified' revision that strips unanswerable questions is an underappreciated mechanical route to a Yes. I weight the trajectory-based Monte Carlo (35–40%) more heavily than the thin, volatile $18K Polymarket quote (56%, range 45–71.5%), landing modestly below both prior forecasts at 40% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 19$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' listed on the official agi.safe.ai leaderboard as of now, and which model holds it?
  2. Does the agi.safe.ai leaderboard report tool-augmented/search-enabled scores (e.g., Gemini 3 Pro with search+code ~45%, GPT-5 with tools, Grok 4 Heavy ~44%) or only no-tools scores?
  3. What has been the month-over-month rate of improvement in top HLE score from January 2025 through late 2025/early 2026 (e.g., ~3% → ~26% → ~38-45%)?
  4. What frontier model releases are expected in 2026 (GPT-5.5/GPT-6, Gemini 3.5/4, Claude 5, Grok 5) and have any labs pre-announced or leaked HLE scores near or above 60%?
  5. Are there known ceiling effects on HLE (label noise, ~10-15% erroneous questions) that would cap achievable accuracy below 60%?
  6. What do the sibling Polymarket threshold markets (e.g., 50%, 60%, 70% by end of 2026) imply about the crowd's distribution of the 2026 max HLE score?
Planner reasoning
This is a Polymarket question about whether any model reaches ≥60% on Humanity's Last Exam (per the agi.safe.ai leaderboard) during 2026. The key drivers are the current leaderboard SOTA, the rate of progress over the past 12-18 months, and whether tool-augmented/agentic scores are listed on the official leaderboard (which would sharply raise the odds). I'll anchor on the Polymarket price, gather leaderboard/news evidence, check sibling threshold markets for an implied distribution, and do a simple trend extrapolation.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved on Humanity’s Last Exam in 2026 be 60% or higher?** - Current price (probability): 56.00% - 7-day price change: -7.00% - 30-day price change: +11.00% - Total volume: $18,131 (USD notional) - Price range: 45.00% - 71.50% - Data point
polymarket_related OK 2.8s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 1 markets | keyword 'AI benchmark 2026': 0 markets | keyword 'best AI model': 0 markets
kalshi_related OK 2.7s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'HLE score': ok
claude_news OK 46.5s 11 Here are the key findings on HLE's current state and score progression: **Official leaderboard status (Scale AI / agi.safe.ai)** - The official Scale Labs no-tools leaderboard currently shows **gemini-3-pro-preview at 37.52%** as the top score, followed by gpt-5.4-2026-03-05 (xhigh thinking) at 36
claude_news OK 26.7s 9 Based on research, here are the key findings: - **Current leaderboard status is contested/uncertain due to model-name confusion in secondary sources**: Wikipedia's current table (citing the official CAIS/Scale leaderboard) lists Gemini 3.1 Pro Preview at 46.44% and GPT-5.4 Pro at 44.32% as the to
gdelt_news OK 115.5s 30 GDELT: 30 articles across 2 queries (lookback=120d). "Humanity's Last Exam benchmark score": 15 hits | 'HLE accuracy leaderboard model': 15 hits
wikipedia OK 0.1s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 72.0s 0 ## Key Findings **Trend fitting (raw data, no adjustment):** - Linear extrapolation from 4 data points (Jan '25 ~6% → Apr '25 ~20% → Aug '25 ~25.5% → Nov '25 ~38-45%) projects **78–92%** by Dec 2026 — implausibly aggressive given it assumes the recent steep slope holds unchanged for 14 more months.
3. Evidence Brief Sonnet · 7392 chars
# Current state The official resolution source (agi.safe.ai / Scale AI CAIS leaderboard, no-tools) currently shows top scores in the high-30s to mid-40s% — well short of 60% — with gemini-3-pro-preview / Gemini 3.1 Pro Preview leading (~37.5–46.4% depending on snapshot). Separately, unofficial third-party trackers (BenchLM, Artificial Analysis) report much higher figures (55–65%) for newer models (Claude Opus 5, Claude Fable 5, Muse Spark 1.1), but these appear to reflect different protocols (tools, ensembles, or disputed methodology) and are NOT yet reflected on the official agi.safe.ai leaderboard that governs resolution. # Timeline of key events - 2025-01: HLE launches; GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0% (confirmed, intuitionlabs.ai). - 2025-mid: Top score climbs to ~20–25.5% across successive model releases (confirmed, code_execution trend data). - 2025-07: FutureHouse investigation finds ~30% of text-only chem/bio HLE questions may be erroneous; HLE team partially replicates and plans revisions ("HLE-Verified") (confirmed). - 2025-11: Top no-tools score reaches ~38–45% per aggregated trackers (reported, mixed sourcing). - 2026 (research snapshot, ~Aug 2026): Official Scale/CAIS leaderboard shows gemini-3-pro-preview 37.52%, gpt-5.4 (xhigh) 36.24%, claude-opus-4-7 36.20% (confirmed, labs.scale.com). A later/Wikipedia-cited snapshot shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32% (reported, possible lag/version difference). - 2026-07-09: Meta's Muse Spark 1.1 benchmark claims disputed over conflict-of-interest allegations (reported, digg.com). - 2026-08: Independent trackers (BenchLM, Artificial Analysis) report Claude Opus 5 (64.7%), Claude Mythos 5 (64.5%), Muse Spark 1.1 (62.1%), Claude Fable 5 (55.5%) — discrepant from official leaderboard, likely tool-augmented/ensemble or unverified model names (rumored/unconfirmed). # Event Will the official agi.safe.ai HLE leaderboard show any model reaching ≥60% accuracy by Dec 31, 2026? # Outcomes to forecast - Yes (≥60% achieved) - No (stays below 60%) # Kalshi market anchor No direct kalshi_direct price was returned in this research pull. The matching Polymarket market (same ticker/question) trades at **56% YES**, down 7pts over 7 days but up 11pts over 30 days; range 45–71.5% over 24 days; volume ~$18.1K. This is the best available cross-market consensus proxy and should be treated as the anchor absent a distinct Kalshi quote. # Sub-question answers 1. **Current top official score?** — gemini-3-pro-preview at 37.52% (no-tools) per labs.scale.com; a later Wikipedia-cited snapshot shows Gemini 3.1 Pro Preview at 46.44% [claude_news/Wikipedia]. Discrepancy suggests dataset update lag between sources. 2. **Tools vs no-tools reporting?** — agi.safe.ai's primary leaderboard is no-tools only; a separate "with tools" leaderboard (llm-stats.com) shows Tencent's Hy3 at 53.2%. The question's resolution source (agi.safe.ai) is no-tools-based, so tool-augmented scores likely don't count directly [claude_news]. 3. **Rate of improvement?** — Roughly 2.7-8% (Jan 2025) → ~20% (mid-2025) → ~25.5% (Aug 2025) → ~38-45% (Nov 2025) → ~37-46% official (mid-2026), suggesting deceleration/plateauing in official no-tools scores despite continued frontier releases [code_execution, claude_news]. 4. **2026 frontier releases w/ leaked 60%+ scores?** — GPT-5.4/5.5, Gemini 3/3.1 Pro, Claude Opus 4.6/4.7 all released but official scores remain <47%; unofficial trackers claim Claude Opus 5/Fable 5/Mythos 5 at 55-65%, but these are unverified/possibly non-official-protocol and not yet on agi.safe.ai [claude_news, benchlm.ai]. 5. **Ceiling effects/label noise?** — Yes: FutureHouse found ~30% error rate in text-only chem/bio questions; HLE team acknowledges issue and plans "HLE-Verified" revisions, which could either cap scores (noise floor) or raise them (if revisions remove unanswerable questions) [claude_news]. 6. **Sibling threshold markets?** — Polymarket price of 56% for this exact 60%-threshold market implies moderate-to-lean-yes crowd sentiment, down from a 71.5% high but up 11pts over 30 days — reflecting volatility/uncertainty rather than consensus. # Key facts (high-confidence, factual) 1. [labs.scale.com] Official no-tools leaderboard top score ~37.5% (gemini-3-pro-preview) at latest official snapshot captured. 2. [Wikipedia/CAIS] A more recent-looking snapshot shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%. 3. [intuitionlabs.ai] HLE launched Jan 2025 at single-digit top scores; by ~1 year later, still under 40-46% on official tracking. 4. [claude_news/FutureHouse] ~30% of chem/bio HLE questions may be flawed/erroneous, partially confirmed by HLE team. 5. [benchlm.ai, artificialanalysis.ai] Independent (non-official) trackers report 55-65% scores for newer Claude variants, unconfirmed on official source. # Cross-market signals - Kalshi related: no directly relevant AI-benchmark markets found; unrelated matches (TV release dates, SCOTUS, NBA) returned by keyword search — no arbitrage signal. - Polymarket (same market): 56% YES, volatile (45-71.5% range), 30d uptrend +11pts but 7d downtrend -7pts — reflects genuine market uncertainty, not consensus collapse toward either outcome. - No sportsbook data applicable. # Analyst opinions and speculation - claude_news bottom line: "close call... hinges on which leaderboard/model claims are verified as legitimate" — high methodological uncertainty. - code_execution Monte Carlo model: P(≥60%) ≈ 35-40%, P(≥50%) ≈ 52%, median trajectory ~52% by Dec 2026, with wide 10-90th percentile band (20-76%). - Some commentary flags conflict-of-interest/gaming concerns (Muse Spark dispute) that could taint any claimed high score's legitimacy for official resolution. # Directional lean per outcome - **Yes (≥60%)**: Supported by extremely rapid historical growth curve (single digits→40s in ~18 months) and unofficial trackers already claiming 55-65% scores for newer models; opposed by the fact these are NOT on the official resolution source and official no-tools progress has visibly decelerated (37-46% plateau in 2026), plus known label-noise ceiling effects. - **No (<60%)**: Supported by official leaderboard's clear deceleration (mid-40s%, far from 60%) and structural noise ceiling (~30% flawed questions in some domains) limiting max achievable accuracy; opposed by continued frontier releases through 2026 and steep historical trajectory that could still close the gap in remaining months. # Gaps / unknowns - No confirmed kalshi_direct quote retrieved in this research pass — using Polymarket as the only cross-market anchor. - Unclear whether "HLE-Verified" revisions (correcting flawed questions) will lower or raise achievable ceiling. - Legitimacy/methodology of high unofficial scores (Opus 5, Fable 5, Mythos 5) unverified — could be tool-augmented, ensemble, or inflated/gamed claims. - Exact current live agi.safe.ai figure and update cadence for rest of 2026 not confirmed. # Calibration anchors - Polymarket cross-market price: 56% YES (as of research date). - Quant model estimate: ~35-40% probability of ≥60% by Dec 2026. - Precedent: benchmarks (MMLU, GPQA) have shown rapid saturation once labs specifically target them, but HLE was designed adversarially to resist quick saturation — historically slower climbs post-initial gains are common for expert-curated benchmarks.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.34
Yes 50%
No 50%
The official resolution source (agi.safe.ai / Scale CAIS no-tools leaderboard) shows top scores of only ~37.5-46.4% (Gemini 3.1 Pro Preview), meaning a ~14-point jump would be needed in the remaining months of 2026, and official no-tools progress has visibly decelerated from its 2025 pace (single digits → ~25% → ~38-46%). Against that, unofficial trackers (BenchLM, Artificial Analysis) already report 55-65% for Claude Opus 5/Mythos 5/Muse Spark 1.1, which — if any is verified under the official no-tools protocol and posted — would resolve Yes immediately; however these are disputed (Muse Spark conflict-of-interest allegations) and possibly tool-augmented. Label noise (~30% flawed chem/bio questions per FutureHouse) imposes a real ceiling, though 'HLE-Verified' revisions could cut either way. The cross-market anchor (Polymarket 56% YES, volatile 45-71.5%) sits above the quant Monte Carlo estimate of 35-40%, so I settle near the midpoint at ~50%.
gpt-5.5
0.55
Yes 43%
No 57%
Anchoring on the available cross-market consensus proxy of 56% YES, I move downward because the official agi.safe.ai no-tools leaderboard still appears to be only ~37.5% to 46.4%, which points toward No because the resolved source is materially short of 60%. The tools-vs-no-tools evidence also points toward No because several 55-65% claims appear to come from unofficial or possibly tool/ensemble protocols that may not count for resolution. Historical improvement from single digits to the 40s points toward Yes because frontier models have improved rapidly, but the apparent 2026 deceleration and label-noise concerns point back toward No. The unofficial high scores keep substantial upside risk, but without official verification I put the event below the market anchor.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters anchor heavily on Polymarket's 56% and drift only modestly (50%/43%), but the quant Monte Carlo model embedded in the brief explicitly estimates P(≥60%)≈35-40% with median trajectory ~52% — neither forecaster explains why they weight the volatile, thin-volume ($18.1K) Polymarket price over this more granular internal model, which seems like a stronger fit to the actual evidence (deceleration curve, noise ceiling). 2. Neither forecast grapples with the concrete arithmetic: going from ~37-46% to ≥60% in under 12 months would require a bigger absolute jump than any prior 12-month period in the benchmark's history (the fastest gain was ~25%→~40% over ~6 months, still short of the ~14-23 point jump needed here from the higher baseline) — this reference-class comparison argues for a lower probability than either gave. 3. Both treat the "HLE-Verified" revision as a wildcard but don't consider the scenario where revisions are explicitly designed to remove the ~30% flawed chem/bio questions, which could mechanically inflate scores on remaining "valid" questions well past 60% by definition-change rather than genuine capability gain — a distinct pathway to Yes that's underexplored. 4. Neither addresses the specific risk that if any unofficial 55-65% claim (e.g., Muse Spark 1.1) is eventually validated and posted to the official leaderboard despite conflict-of-interest disputes, resolution could hinge on a contested/political process rather than a clean technical threshold — this adds tail-risk uncertainty on both sides that isn't reflected in either point estimate. 5. Forecast 1's "midpoint" methodology (splitting the difference between Polymarket 56% and quant model 35-40%) is not a principled aggregation — it's an arbitrary average that ignores that the quant model was built specifically on the confirmed leaderboard trajectory data, which should arguably carry more weight than a thin, volatile prediction market. 6. Both forecasts land close to each other (50% vs 43%) despite the brief's own analyst quote calling this "a close call" hinging on verification — neither forecaster explicitly prices in the elevated resolution-criteria ambiguity (no-tools vs. tools, disputed leaderboard entries) as a source of variance that might justify wider uncertainty or a more clearly stated "lean No" given the official source is the sole resolution criterion and sits ~14+ points below threshold.
6. Aggregator Opus · final calibrated forecast
40% (45% confidence)
Yes: 40% No: 60%
Resolution hinges solely on the official agi.safe.ai/Scale CAIS no-tools leaderboard, which sits at ~37.5–46.4% — a required jump of ~14–23 points in roughly the final four months of 2026, larger in absolute terms than any prior comparable window and against a visibly decelerating official curve. The 55–65% figures from BenchLM/Artificial Analysis are the main Yes pathway but are unverified, plausibly tool-augmented or ensemble-based, and in one case (Muse Spark 1.1) actively disputed for conflict of interest; a contested posting adds two-sided tail risk rather than clean upside. Label-noise (~30% flawed chem/bio items) imposes a genuine ceiling, though an 'HLE-Verified' revision that strips unanswerable questions is an underappreciated mechanical route to a Yes. I weight the trajectory-based Monte Carlo (35–40%) more heavily than the thin, volatile $18K Polymarket quote (56%, range 45–71.5%), landing modestly below both prior forecasts at 40% Yes.
7. Entry Decision Opus Trader
Side
NO
Position Size
$1,100
Contracts
2444
Trader Confidence
44%
Trade Rationale
This is a direction-disagreement setup: the bot sits at 40% while the market prices 56%, opposite sides of 50%, which is the historically cleanest edge type. The core arithmetic is concrete and checkable — the official CAIS no-tools leaderboard is at ~37.5–46.4%, so YES requires a 14–23 point jump in ~4.5 months against a visibly decelerating curve, larger in absolute terms than any prior comparable window. The main YES pathways (BenchLM/Artificial Analysis 55–65% claims, one disputed for conflict of interest, plus a hypothetical 'HLE-Verified' rescoring) are unverified and may not be leaderboard-eligible, so they are tail routes rather than base case. Forecaster confidence is a middling 0.45 with a 7pp ensemble spread and one member at essentially a coin flip (50%), and the Devil's Advocate correctly flags the definition-change route to YES as underweighted — so I take the trade but keep it near baseline rather than large.
Allocation Logic
$1100 — slightly above the $1000 baseline for a 16pp direction-disagreement edge, but held back from $1500+ because of the genuine two-sided tail risk (contested leaderboard postings, HLE-Verified rescoring), only middling confidence, a thin $18K book, and existing correlated AI-benchmark exposure (Gemini Pro HLE NO, Anthropic model NO).
Entry price: $0.45
Current: $0.30
Status: OPEN
P&L: -$378.89
Pipeline Timing
Total pipeline time: 222.3s
Per-tool research timings shown in the Research section above.