# Current state
As of Aug 2026, the **official CAIS/Scale AI leaderboard (agi.safe.ai / labs.scale.com)** — the market's stated resolution source — tops out at **~46.4%** (Gemini 3.1 Pro preview), still below the 60% threshold and still being actively updated. Independent trackers (Artificial Analysis, BenchLM, PricePerToken) using different protocols (adaptive reasoning, "max effort," or tool-assisted runs) report Claude-family models scoring 53–65%, but it is unconfirmed whether these match the official site's no-tools "HLE Accuracy" metric used for resolution.
# Timeline of key events
- 2025-01: HLE released; baseline scores near-zero (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0%) — confirmed (IntuitionLabs/Wikipedia).
- 2025-mid: o3-based Deep Research ~26.6% — reported.
- 2025-11-18: Gemini 3 Pro 37.5% (no tools); Gemini 3 Deep Think 41.0% (no tools) — confirmed (Google blog).
- 2025-11 (pre-Gemini 3): GPT-5 Pro top score 31.6% — reported (AOL/press).
- 2026-02-12: Gemini 3 Deep Think updated to 48.4% (no tools), official Google claim — confirmed (Google blog).
- 2026-06: Claude Fable 5 launches, 53% HLE — reported (Artificial Analysis).
- 2026-07-09: Meta's Muse Spark 1.1 HLE claims disputed as conflict-of-interest — reported (Digg).
- 2026-07-24: Claude Opus 5 — 56.3% no-tools / 64.7% with tools — reported (MarkTechPost).
- 2026-08 (~mid): Official Scale/CAIS leaderboard shows Gemini 3.1 Pro (46.44%) and GPT-5.4 Pro (44.32%) as top two — reported (claude_news synthesis).
- 2026-08-11: PricePerToken leaderboard: Fable 5 55.5%, Opus 5 54.9%, GPT-5.6 Sol 49.5% — reported.
- 2026-08-22: BenchLM (third-party, tool-inclusive) shows Opus 5 64.7%, Mythos 5 64.5%, Muse Spark 1.1 62.1% — reported, methodology disputed.
# Event
Will the highest score on any model on the official Humanity's Last Exam leaderboard (agi.safe.ai) reach ≥60% "HLE Accuracy" by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No direct kalshi_direct quote was returned for this ticker; the ticker string matches a Polymarket condition ID and the only direct price data available is from **polymarket_direct**: current YES/"≥60%" price **72.5%**, up sharply from 45% a month ago (+27.5% in 30d, +11.5% in 7d), volume ~$23k. Treat this as the closest available consensus anchor pending confirmed Kalshi data — it has moved strongly toward Yes recently, likely driven by Claude Opus 5/BenchLM headlines.
# Sub-question answers
1. **Current highest official HLE score** — Official Scale/CAIS leaderboard tops at ~46.44% (Gemini 3.1 Pro preview, thinking-high), with GPT-5.4 Pro at 44.32%, as of Aug 2026 (claude_news).
2. **Month-by-month SOTA progression** — Roughly 9%→41% (Jan–Nov 2025, ~2.7–3.3 pts/mo), then a marked slowdown to ~1.3 pts/mo through early 2026 (48.4% Feb 2026), before third-party trackers jump to 53–65% by mid-2026 via Claude Fable 5/Opus 5 (code_execution Monte Carlo synthesis).
3. **Tools vs. no-tools on leaderboard** — Official leaderboard historically reports no-tools scores; the 60%+ figures (Opus 5 64.7%) are explicitly tool-augmented per MarkTechPost/BenchLM, while no-tools scores remain mid-50s (Artificial Analysis 55.5%). Ambiguous whether official site would ever list tool-augmented numbers as "HLE Accuracy."
4. **Expected frontier releases before Dec 2026** — GPT-5.5/5.6, Claude Fable 5/Opus 5/Mythos 5, Gemini 3.1 Pro/Deep Think updates, Meta Muse Spark all already released by mid-2026 per timeline; further GPT-6/Gemini 4/Grok 5 releases plausible before year-end (code_execution notes 2-3 more major releases expected).
5. **Is agi.safe.ai still active?** — Yes, confirmed still updated with new entries as of Aug 2026 (claude_news), though it lags behind third-party trackers.
6. **Adjacent markets** — No other Kalshi HLE-threshold markets found (kalshi_related returned irrelevant matches). Polymarket price for this same question is 72.5% and rising fast.
# Key facts (high-confidence, factual)
1. [Google blog] Gemini 3 Deep Think reached 48.4% (no tools) on HLE by Feb 2026 — official lab claim.
2. [claude_news/Scale] Official CAIS/Scale leaderboard top score ~46.44% (Gemini 3.1 Pro) as of Aug 2026 — below 60%.
3. [MarkTechPost] Claude Opus 5: 56.3% no-tools, 64.7% with tools (Jul 2026).
4. [Artificial Analysis] Independent no-tools tracker's current max is 55.5% (Claude Fable 5, Aug 2026) — still under 60%.
5. [Wikipedia] HLE co-created by CAIS and Scale AI; agi.safe.ai is CAIS's official site.
# Cross-market signals
- Kalshi related: no directly comparable HLE-threshold markets found.
- Polymarket: same-question price 72.5%, up from 45% a month ago — strong recent Yes momentum.
- Sportsbook implied: n/a.
# Analyst opinions and speculation
- code_execution Monte Carlo: blended P(≥60%) ≈ 45–55%, but highly scenario-dependent (18% conservative vs 97% optimistic) depending on whether Aug–Nov 2025 slowdown reflects a real ceiling near 55–65% or a temporary lull.
- claude_news: resolution "likely hinges on which leaderboard/methodology is treated as authoritative" — official site lags well behind tool-assisted third-party claims.
- Digg commentary flags conflict-of-interest concerns about self-reported benchmark scores (Meta Muse Spark), raising doubt about inflated third-party numbers generally.
# Directional lean per outcome
- **Yes**: Rapid 2025 trajectory, multiple 50%+ no-tools claims by mid-2026 (Fable 5, Opus 5), momentum in adjacent Polymarket price (72.5%), several more frontier releases expected before Dec 2026.
- **No**: The specific resolution source (official agi.safe.ai/Scale leaderboard) sits well below 60% (~46%) as of Aug 2026 and has shown decelerating no-tools progress; ambiguity whether tool-assisted 60%+ scores would even count; strict resolution rules favor the stricter, lagging official metric.
# Gaps / unknowns
- No confirmed live snapshot of the exact agi.safe.ai number at time of forecast (only ~Aug 2026 secondhand reports).
- Unclear whether official leaderboard will ever adopt tool-augmented scoring before Dec 2026 close.
- True Kalshi YES price for this ticker not directly retrieved (only Polymarket data available, though ticker matches).
- Four months (Sep–Dec 2026) of further progress unaccounted for; could close gap from 46%→60% if trend continues at even 2-3 pts/mo.
# Calibration anchors
- Polymarket price for same question: 72.5% (rising fast).
- Official/resolution-relevant score trajectory: ~9% (Jan 2025) → ~46% (Aug 2026), ~2 pts/mo average over 19 months; needs another ~14 pts in ~4 months to hit 60% via official metric — faster than trailing-12-month average pace, but consistent with historical jump-driven accelerations at major model launches.