# Current state
The official resolution source (agi.safe.ai / Scale AI leaderboard) shows a top HLE Accuracy of **~46.44%** (Gemini 3.1 Pro Preview, as of 2026-08-16), still well below the 60% threshold. Third-party trackers (Artificial Analysis, llm-stats.com) report unofficial/self-reported scores of 55–65% for newer Claude "Opus 5"/"Fable 5" variants, but these are NOT reflected on the official leaderboard and are explicitly flagged as unverified self-reported results.
# Timeline of key events
- 2025-01: HLE launched; SOTA models scored single digits (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (Wikipedia).
- 2025-06: Grok-4 (with tool use) reported ~38–45% — reported (mixed sourcing on tools vs. no-tools).
- 2025-08: GPT-5 reported ~25–42% depending on configuration — reported.
- 2025-11: Gemini 3.0 Pro reached ~37–45%, official "highest recorded" cited as ~38.3% — reported.
- 2026 (early): Official leaderboard scores clustered ~38–40% — reported.
- 2026-07-09: Meta's "Muse Spark 1.1" benchmark claims disputed amid conflict-of-interest allegations — reported/rumored.
- 2026-07-10: GPT-5.6 Sol/Terra/Luna launched — confirmed release, HLE scores not specified in research.
- 2026-08-16: Official leaderboard snapshot: Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Claude Opus 4.6 (Thinking) 34.44% — confirmed (Wikipedia/IntuitionLabs).
- 2026-08 (ongoing): Artificial Analysis reports Claude Fable 5 (55.5%) and Claude Opus 5 (54.9%) — reported, unofficial tracker, not on agi.safe.ai.
- 2026-08 (ongoing): llm-stats.com reports Claude Opus 5 at 64.7%, but flags "0 verified / 101 self-reported" results — rumored/unverified.
# Event
Will the highest score on the official Humanity's Last Exam leaderboard (agi.safe.ai) reach ≥60% accuracy by Dec 31, 2026?
# Outcomes to forecast
- Yes (≥60% achieved by any model on official leaderboard)
- No (highest score stays below 60%)
# Kalshi market anchor
No direct Kalshi YES price was returned by kalshi_direct; only a cross-listed Polymarket price for the identical event is available: **73.5% YES** (as of latest snapshot), up from 45% a month ago (+12.5% 30-day trend, flat 7-day). Volume is thin ($23,866 total over 39 days), suggesting a fairly illiquid, still-forming consensus. This should be treated as the best available proxy anchor given the missing native Kalshi quote.
# Sub-question answers
1. **Current highest official HLE score** — Gemini 3.1 Pro Preview at 46.44% leads the official Scale AI/agi.safe.ai leaderboard as of 2026-08-16 [claude_news/Wikipedia/IntuitionLabs].
2. **Trajectory of top scores** — Roughly 3–9% (Jan 2025) → ~25–45% (mid/late 2025) → ~38–46% (2026 to date). Growth has been rapid but decelerating; no model has yet crossed 50% on the *official* board [claude_news, code_execution].
3. **Recent models near/above 60%** — Not on the official leaderboard. Third-party sites (Artificial Analysis: Claude Fable 5 ~55.5%, Opus 5 ~54.9%; llm-stats.com: Opus 5 ~64.7%) report higher figures, but these are self-reported/unverified and use possibly different methodology (tool-use, adaptive reasoning) not confirmed to match "HLE Accuracy" as officially defined [claude_news].
4. **Leaderboard update cadence** — Updates continuously as new frontier models are verified, but with lag and a private held-out set to detect overfitting; it does not automatically absorb third-party/self-reported tool-augmented results [claude_news].
5. **Expected 2026 releases** — GPT-5.6 (Sol/Terra/Luna), Gemini 3.7 Flash, DeepSeek V4-Pro already released or imminent per GDELT; none yet confirmed to push official HLE above 50–55% [gdelt_news].
6. **Adjacent market pricing** — No native Kalshi ≥50%/≥70% threshold markets found; Polymarket price for this exact 60% threshold is 73.5%, implying markets lean Yes but with recent upward momentum, not settled certainty.
# Key facts (high-confidence, factual)
1. [Wikipedia/claude_news] Official leaderboard top score ≈46.44% (Gemini 3.1 Pro Preview) as of 2026-08-16.
2. [claude_news] HLE launched Jan 2025 with SOTA scores in single digits.
3. [claude_news] llm-stats.com explicitly notes 0 of 101 evaluated results are "verified" — all self-reported.
4. [polymarket_direct] Polymarket price for this exact question: 73.5% YES, up sharply (+12.5pts) over 30 days.
5. [code_execution] Trend-extrapolation models disagree widely: linear fits imply 67–95%; logistic (saturating) fits imply 50–68%, straddling the 60% line.
# Cross-market signals
- Kalshi related: No directly comparable AI-benchmark markets found; unrelated hits (TV show release, SCOTUS case) not informative.
- Polymarket: Same-event price 73.5% YES, with strong 30-day upward momentum (45%→73.5%), thin volume (~$24k) — moderate confidence signal.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Mid-2025 analysts predicted "50% by end of 2025" — did not materialize on official board [claude_news].
- Some commentary suggests official/third-party score gap (contamination-control private set, verification lag) may cause official scores to lag true frontier capability by many months, which could delay a Yes resolution even if capability effectively exceeds 60%.
- Code-execution quantitative synthesis leans mildly Yes (~55-60% probability) but flags extreme sensitivity to model choice/tools-inclusion.
# Directional lean per outcome
- **Yes**: Rapid historical growth curve (9%→46% in ~20 months); multiar third-party trackers already claim 55–65%; frontier labs (GPT-5.6, Gemini 3.7, Claude Opus 5) continuing rapid iteration; Polymarket pricing 73.5%.
- **No**: Official leaderboard still 14+ points below threshold with only ~4 months of 2026 lapsed at last snapshot; unofficial high scores are unverified/self-reported and may not count; logistic/saturation models suggest diminishing returns as HLE was designed to resist saturation; benchmark stewards periodically raise difficulty via private held-out sets.
# Gaps / unknowns
- No direct Kalshi native quote retrieved — only Polymarket cross-price used as proxy anchor.
- Unclear whether official leaderboard will incorporate/verify the high third-party scores (Claude Opus 5, 54-65%) before Dec 2026.
- No confirmed frontier-lab roadmap data for Q4 2026 releases.
# Calibration anchors
- Cross-market anchor (Polymarket, same event): 73.5% YES, rising.
- Official leaderboard current top score: 46.44%, needs +13.5 points in ~4 months to resolve Yes — a steep but not unprecedented jump given historical ~10-15pt gains per 3-6 months.
- Base rate precedent: HLE scores rose ~35-40 points in 20 months; achieving another 13+ points in 4 months is plausible but requires continued rapid pace.