# Current state
The market resolves on the OFFICIAL agi.safe.ai (Scale AI/CAIS) Humanity's Last Exam leaderboard. That leaderboard, as of the most recent snapshot (~Aug 2026), shows the best-ranked Claude model (Opus 4.7) at only ~36% (no-tools) — well below 60%. However, Anthropic's own July 2026 Opus 5 launch reports a no-tools HLE score of 56.3% and a with-tools score of 64.7%, figures not yet reflected on the official leaderboard snapshot cited. This creates a live reconciliation gap between self-reported Anthropic numbers and the lagging official leaderboard that will determine resolution.
# Timeline of key events
- 2025-01: HLE launches; frontier models score single digits (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (intuitionlabs.ai).
- 2025-11-18: Gemini 3 Pro sets record 37.4% (no-tools), surpassing GPT-5 Pro's 31.64% — confirmed (TechCrunch).
- 2025-11 (Opus 4.5 launch): Claude trails Gemini 3 Pro by ~7pts no-tools, ~2pts with search — confirmed (Vellum).
- 2026-02: Opus 4.6 system card claims Anthropic "leads all frontier models" on HLE; third-party recompute: 40.0% no-tools (vs GPT-5.2's 50.0%), 53.1% with tools — reported, self-serving framing flagged.
- 2026-02-12: Gemini 3 Deep Think sets new SOTA, 48.4% no-tools — confirmed (MarkTechPost).
- 2026-05-28/29: Claude Opus 4.8 launched (dynamic workflows) — confirmed (moneycontrol, mashable); benchmark specifics not detailed.
- 2026-07-24: Claude Opus 5 launches; Anthropic reports 56.3% no-tools / 64.7% with tools on HLE — reported by Anthropic + independently corroborated (aireleasetracker.com, benchlm.ai).
- 2026-08 (aggregator snapshots): Artificial Analysis/pricepertoken list "Claude Fable 5" at 55.5% and Opus 5 at 54.9%; felloai.com cites Fable 5 at 53.3% — conflicting, unverified model-naming caveat (Fable/Mythos not on Anthropic's standard Opus/Sonnet/Haiku scheme per Wikipedia; Wikipedia does confirm a 2026 "Mythos"/"Fable" release exists).
- 2026-08 (official Scale/agi.safe.ai snapshot): top Claude entry still listed as Opus 4.7 at ~36.2%, i.e., official leaderboard appears NOT yet updated with Opus 5 — reported, key lag/omission risk.
# Event
Will any Anthropic Claude model reach ≥60% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct price returned for this specific ticker; nearest cross-market anchor is Polymarket at 50.5% YES (down 2pts/7d, up 5.5pts/30d), range 17–60.5% over 20 days, ~$19.3K volume — treat as the consensus to beat.
# Sub-question answers
1. **Current highest Claude HLE figure on official leaderboard** — ~36.2% (Claude Opus 4.7, no-tools) per Scale/agi.safe.ai Aug 2026 snapshot (claude_news). Self-reported Opus 5 (Jul 2026) claims 56.3% no-tools but not yet visible on the official leaderboard.
2. **SOTA across labs / tools vs no-tools** — Official leaderboard SOTA no-tools ~46–48% (Gemini 3.1 Pro Preview, GPT-5.4 Pro); Gemini 3 Deep Think claimed 48.4%. Leaderboard is primarily no-tools; "with tools" scores exist but are typically self-reported in model cards, not the primary leaderboard column.
3. **Rate of increase / extrapolation** — Naive Fermi/code_execution fit (using data only through Nov 2025) implies Claude ≈19.7–57.2% by Dec 2026 (mean ~34%, P(≥60%)≈10-25%). This is now stale: actual Jul 2026 Opus 5 self-reported score (56.3%) already exceeds that mean, showing real trajectory outpaced the earlier linear/logistic fits.
4. **Leaderboard update lag** — Evidence of a real lag: Opus 5 (Jul 2026) not reflected in Aug 2026 official snapshot showing only Opus 4.7. Omission risk is confirmed and material for year-end resolution timing.
5. **2026 release cadence / targets** — Confirmed cadence: Opus 4.6 (Feb), Opus 4.8 (May), Sonnet 5 (mid-2026), Opus 5 (Jul 24). Wikipedia confirms additional restricted "Mythos" and public "Fable" variants released in 2026. No explicit numeric HLE target stated by Anthropic beyond claiming benchmark leadership.
6. **Parallel markets** — No other Kalshi/Polymarket HLE threshold markets found (0 keyword matches for "HLE"/"Humanity's Last Exam"); only this Polymarket contract exists as a direct comparator.
# Key facts (high-confidence, factual)
1. [claude_news/Scale AI] Official leaderboard Aug 2026: top Claude entry (Opus 4.7) ≈36.2%, below Gemini/GPT SOTA (~44-46%).
2. [claude_news/Anthropic] Opus 5 (Jul 24, 2026) self-reported: 56.3% no-tools, 64.7% with tools.
3. [TechCrunch/MarkTechPost] Non-Anthropic SOTA progressed 37.4%→48.4% no-tools between Nov 2025–Feb 2026.
4. [Wikipedia] Anthropic released additional 2026 model tiers "Mythos" (restricted) and "Fable" (public), beyond standard Opus/Sonnet/Haiku naming — corroborates aggregator "Fable 5" mentions.
5. [Polymarket] Current YES 50.5%, range 17–60.5%, trending up over 30d.
# Cross-market signals
- Kalshi related: no direct HLE market; tangential Anthropic markets (IPO race 92%, sector classification 85%) show no benchmark-relevant signal.
- Polymarket: 50.5% YES, coin-flip pricing, recent modest uptrend (+5.5% 30d) consistent with Opus 5's July release news.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- claude_news synthesis: "if with-tools counts, threshold arguably already cleared; if only no-tools counts, Claude is just under 60% but climbing rapidly."
- Digg/commentators flag conflict-of-interest concerns about self-reported benchmark claims (Meta Muse Spark case) — generalizable caution for Anthropic's own Opus 5 figures pending official leaderboard confirmation.
- code_execution Fermi model (using pre-Opus5 data) is stale/underestimates given Opus 5's actual reported jump to 56.3%.
# Directional lean per outcome
- **Yes**: Opus 5 no-tools (56.3%) is very close to threshold; with-tools (64.7%) already exceeds it; steep recent trajectory (40%→56.3% in ~5 months) plus ~5 more months and likely further Opus/Sonnet iterations before Dec 2026 close.
- **No**: Official agi.safe.ai leaderboard still shows only ~36%, with confirmed update lag/omission of Opus 5; resolution explicitly keys on "HLE Accuracy" label on official source, likely the no-tools column, and leaderboard methodology updates (grader revisions) have sometimes lowered reported scores; unverified aggregator model names ("Fable 5") add noise/risk of overstatement.
# Gaps / unknowns
- Whether official leaderboard will list a "with tools" HLE Accuracy figure that could trigger Yes at 64.7%.
- Timing/likelihood of agi.safe.ai actually updating with Opus 5 (or later models) before Dec 31, 2026 cutoff.
- True status/naming of "Claude Fable 5"/"Mythos" — possibly same models as Opus 5 variants, unclear if independently verified on official leaderboard.
# Calibration anchors
- Polymarket YES 50.5% (primary cross-market anchor).
- Precedent: HLE SOTA has moved from single digits (Jan 2025) to ~48% (Feb 2026) to Anthropic's self-reported 56%+ (Jul 2026) — a fast-moving benchmark where large multi-month jumps are historically plausible.