# Current state
The official resolution source (agi.safe.ai / Scale AI CAIS leaderboard, no-tools) currently shows top scores in the high-30s to mid-40s% — well short of 60% — with gemini-3-pro-preview / Gemini 3.1 Pro Preview leading (~37.5–46.4% depending on snapshot). Separately, unofficial third-party trackers (BenchLM, Artificial Analysis) report much higher figures (55–65%) for newer models (Claude Opus 5, Claude Fable 5, Muse Spark 1.1), but these appear to reflect different protocols (tools, ensembles, or disputed methodology) and are NOT yet reflected on the official agi.safe.ai leaderboard that governs resolution.
# Timeline of key events
- 2025-01: HLE launches; GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0% (confirmed, intuitionlabs.ai).
- 2025-mid: Top score climbs to ~20–25.5% across successive model releases (confirmed, code_execution trend data).
- 2025-07: FutureHouse investigation finds ~30% of text-only chem/bio HLE questions may be erroneous; HLE team partially replicates and plans revisions ("HLE-Verified") (confirmed).
- 2025-11: Top no-tools score reaches ~38–45% per aggregated trackers (reported, mixed sourcing).
- 2026 (research snapshot, ~Aug 2026): Official Scale/CAIS leaderboard shows gemini-3-pro-preview 37.52%, gpt-5.4 (xhigh) 36.24%, claude-opus-4-7 36.20% (confirmed, labs.scale.com). A later/Wikipedia-cited snapshot shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32% (reported, possible lag/version difference).
- 2026-07-09: Meta's Muse Spark 1.1 benchmark claims disputed over conflict-of-interest allegations (reported, digg.com).
- 2026-08: Independent trackers (BenchLM, Artificial Analysis) report Claude Opus 5 (64.7%), Claude Mythos 5 (64.5%), Muse Spark 1.1 (62.1%), Claude Fable 5 (55.5%) — discrepant from official leaderboard, likely tool-augmented/ensemble or unverified model names (rumored/unconfirmed).
# Event
Will the official agi.safe.ai HLE leaderboard show any model reaching ≥60% accuracy by Dec 31, 2026?
# Outcomes to forecast
- Yes (≥60% achieved)
- No (stays below 60%)
# Kalshi market anchor
No direct kalshi_direct price was returned in this research pull. The matching Polymarket market (same ticker/question) trades at **56% YES**, down 7pts over 7 days but up 11pts over 30 days; range 45–71.5% over 24 days; volume ~$18.1K. This is the best available cross-market consensus proxy and should be treated as the anchor absent a distinct Kalshi quote.
# Sub-question answers
1. **Current top official score?** — gemini-3-pro-preview at 37.52% (no-tools) per labs.scale.com; a later Wikipedia-cited snapshot shows Gemini 3.1 Pro Preview at 46.44% [claude_news/Wikipedia]. Discrepancy suggests dataset update lag between sources.
2. **Tools vs no-tools reporting?** — agi.safe.ai's primary leaderboard is no-tools only; a separate "with tools" leaderboard (llm-stats.com) shows Tencent's Hy3 at 53.2%. The question's resolution source (agi.safe.ai) is no-tools-based, so tool-augmented scores likely don't count directly [claude_news].
3. **Rate of improvement?** — Roughly 2.7-8% (Jan 2025) → ~20% (mid-2025) → ~25.5% (Aug 2025) → ~38-45% (Nov 2025) → ~37-46% official (mid-2026), suggesting deceleration/plateauing in official no-tools scores despite continued frontier releases [code_execution, claude_news].
4. **2026 frontier releases w/ leaked 60%+ scores?** — GPT-5.4/5.5, Gemini 3/3.1 Pro, Claude Opus 4.6/4.7 all released but official scores remain <47%; unofficial trackers claim Claude Opus 5/Fable 5/Mythos 5 at 55-65%, but these are unverified/possibly non-official-protocol and not yet on agi.safe.ai [claude_news, benchlm.ai].
5. **Ceiling effects/label noise?** — Yes: FutureHouse found ~30% error rate in text-only chem/bio questions; HLE team acknowledges issue and plans "HLE-Verified" revisions, which could either cap scores (noise floor) or raise them (if revisions remove unanswerable questions) [claude_news].
6. **Sibling threshold markets?** — Polymarket price of 56% for this exact 60%-threshold market implies moderate-to-lean-yes crowd sentiment, down from a 71.5% high but up 11pts over 30 days — reflecting volatility/uncertainty rather than consensus.
# Key facts (high-confidence, factual)
1. [labs.scale.com] Official no-tools leaderboard top score ~37.5% (gemini-3-pro-preview) at latest official snapshot captured.
2. [Wikipedia/CAIS] A more recent-looking snapshot shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%.
3. [intuitionlabs.ai] HLE launched Jan 2025 at single-digit top scores; by ~1 year later, still under 40-46% on official tracking.
4. [claude_news/FutureHouse] ~30% of chem/bio HLE questions may be flawed/erroneous, partially confirmed by HLE team.
5. [benchlm.ai, artificialanalysis.ai] Independent (non-official) trackers report 55-65% scores for newer Claude variants, unconfirmed on official source.
# Cross-market signals
- Kalshi related: no directly relevant AI-benchmark markets found; unrelated matches (TV release dates, SCOTUS, NBA) returned by keyword search — no arbitrage signal.
- Polymarket (same market): 56% YES, volatile (45-71.5% range), 30d uptrend +11pts but 7d downtrend -7pts — reflects genuine market uncertainty, not consensus collapse toward either outcome.
- No sportsbook data applicable.
# Analyst opinions and speculation
- claude_news bottom line: "close call... hinges on which leaderboard/model claims are verified as legitimate" — high methodological uncertainty.
- code_execution Monte Carlo model: P(≥60%) ≈ 35-40%, P(≥50%) ≈ 52%, median trajectory ~52% by Dec 2026, with wide 10-90th percentile band (20-76%).
- Some commentary flags conflict-of-interest/gaming concerns (Muse Spark dispute) that could taint any claimed high score's legitimacy for official resolution.
# Directional lean per outcome
- **Yes (≥60%)**: Supported by extremely rapid historical growth curve (single digits→40s in ~18 months) and unofficial trackers already claiming 55-65% scores for newer models; opposed by the fact these are NOT on the official resolution source and official no-tools progress has visibly decelerated (37-46% plateau in 2026), plus known label-noise ceiling effects.
- **No (<60%)**: Supported by official leaderboard's clear deceleration (mid-40s%, far from 60%) and structural noise ceiling (~30% flawed questions in some domains) limiting max achievable accuracy; opposed by continued frontier releases through 2026 and steep historical trajectory that could still close the gap in remaining months.
# Gaps / unknowns
- No confirmed kalshi_direct quote retrieved in this research pass — using Polymarket as the only cross-market anchor.
- Unclear whether "HLE-Verified" revisions (correcting flawed questions) will lower or raise achievable ceiling.
- Legitimacy/methodology of high unofficial scores (Opus 5, Fable 5, Mythos 5) unverified — could be tool-augmented, ensemble, or inflated/gamed claims.
- Exact current live agi.safe.ai figure and update cadence for rest of 2026 not confirmed.
# Calibration anchors
- Polymarket cross-market price: 56% YES (as of research date).
- Quant model estimate: ~35-40% probability of ≥60% by Dec 2026.
- Precedent: benchmarks (MMLU, GPQA) have shown rapid saturation once labs specifically target them, but HLE was designed adversarially to resist quick saturation — historically slower climbs post-initial gains are common for expert-curated benchmarks.