# Current state
Anthropic's most recent flagship, Claude Opus 5 (released 2026-07-24), reportedly scores 56.3% on HLE without tools and 64.7% with tools (self-reported/aggregator-corroborated). However, the *official* resolution source (agi.safe.ai, which mirrors the Scale AI/CAIS leaderboard) still shows Claude's best confirmed no-tools score in the 34–46% range as of August 2026 (latest listed model: claude-opus-4-7 at 36.20%), suggesting the official leaderboard lags newer releases and/or uses a stricter/different scoring methodology than third-party trackers. Whether the market resolves YES hinges on whether the official leaderboard eventually posts a Claude no-tools score ≥60% before 2026-12-31.
# Timeline of key events
- 2025-01: HLE launches; SOTA models score only 3–9% (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — confirmed (Wikipedia/claude_news).
- 2025-02: Claude 3.7 Sonnet ~8.8% — reported.
- 2025-05: Claude 4 Opus (extended thinking) ~23.8%; frontier (any model) ~20.3% — reported.
- 2025-08: Frontier best (any lab) ~25.2–25.4% (Grok 4/GPT-5); Claude Opus 4.1 regresses to ~18.5% — reported (plateau signal).
- 2025-12: Claude Opus 4.5 no-tools score trails Gemini 3 Pro by ~7pp (implying high-20s/low-30s%) — reported (Vellum).
- 2026-02/03: Official Scale AI leaderboard: Gemini 3 Pro 37.5%, Claude Opus 4.6 Thinking Max 34.4%, GPT-5 Pro 31.6% — confirmed (IntuitionLabs).
- 2026-05: Claude Opus 4.8 released; mid-to-high 30s HLE range for Anthropic's top models per trackers — reported.
- 2026-06: "Claude Fable 5" reported at 55.5% (no-tools, third-party) — reported, model naming unconfirmed as official Anthropic branding at time of report.
- 2026-07-24: Claude Opus 5 launched; HLE 56.3% no-tools / 64.7% with tools — reported (Anthropic/aggregators), corroborated across sources.
- 2026-08 (early): Official Scale AI leaderboard still tops out with claude-opus-4-7 at 36.20% (no-tools); text-only variant shows claude-opus-4-6-thinking-max as best Claude entry — confirmed, indicating leaderboard lag vs. Opus 5 release.
- 2026-08-11/19: Third-party trackers (Artificial Analysis, PricePerToken, BenchLM) list Claude Fable 5 (55.5%) and Claude Opus 5 (54.9–64.7% depending on protocol) as top HLE performers overall — reported, methodology/naming caveats apply.
# Event
Will any Anthropic Claude model reach ≥60% HLE accuracy (per agi.safe.ai / Scale AI official leaderboard) by 2026-12-31?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No direct Kalshi price returned for this ticker; Polymarket (same event) shows current YES price 60% (down 4% over 7 days, up 15% over 30 days), volume $26.9K, range 17%–64.5% over 29 days — treat as primary cross-market anchor given absence of Kalshi-native price in the tool output.
# Sub-question answers
1. **Highest official Claude HLE score & protocol** — Official Scale AI leaderboard (Aug 2026): claude-opus-4-7 at 36.20% (no-tools); earlier Feb/Mar 2026 snapshot had Claude Opus 4.6 Thinking Max at 34.4%. All confirmed official figures are no-tools/closed-book. [claude_news/Scale AI]
2. **Best score by any model, gap to Claude** — Gemini 3.1 Pro Preview leads officially at ~46–47% (Aug 2026); GPT-5.4-pro close behind at ~44–45%. Claude trails by ~8–10pp on the official leaderboard. [claude_news]
3. **Rate of improvement** — Frontier rose 9%→20%→25% (Jan–Aug 2025), then plateaued/regressed (Opus 4.1 fell vs Opus 4); by early-mid 2026 official frontier only reached ~37–46%. Growth is decelerating, not exponential. [code_execution, claude_news]
4. **Leaderboard update cadence** — Scale AI's official leaderboard appears to lag new releases by weeks-to-months: Opus 5 (released 2026-07-24, self-reported 56.3% no-tools) is not yet reflected in the Aug 2026 official snapshot, which still shows opus-4-7 at 36.2%. [claude_news, inferred]
5. **Expected 2026 Claude releases** — Confirmed released in 2026: Opus 4.6/4.7/4.8, Claude Mythos, Claude Fable, Claude Sonnet 5, and Claude Opus 5 (2026-07-24). Self-reported/aggregator HLE for Opus 5: 56.3% no-tools, 64.7% with tools. Naming/versioning beyond Opus 5 unconfirmed for H2 2026. [claude_news, Wikipedia]
6. **Sibling markets / implied distribution** — No true HLE sibling markets found on Polymarket or Kalshi (0 matches); only loosely related Kalshi markets (Anthropic IPO odds, US equity stake market) exist, offering no direct calibration. [polymarket_related, kalshi_related]
# Key facts (high-confidence, factual)
1. [claude_news/Scale AI] Official leaderboard best Claude (no-tools) ≈36–46% as of mid-2026, below 60%.
2. [claude_news] Anthropic Opus 5 (2026-07-24) self-reported/aggregator no-tools HLE = 56.3%; with-tools = 64.7%.
3. [code_execution] Logistic-fit extrapolation of official trend data implies ~20% ceiling by end-2026 under status-quo growth; naive exponential fit is invalid (unbounded).
4. [Wikipedia] HLE is a 2,500-question benchmark; no mention of Claude exceeding 60% in encyclopedic record as of last update.
# Cross-market signals
- Kalshi related: No direct HLE/Claude benchmark markets found; adjacent Anthropic markets (IPO odds 93%, US stake market 16%) unrelated to capability trajectory.
- Polymarket: Same-event price at 60% YES, volatile (17%–64.5% range), recent 30-day rise (+15%) suggests momentum toward YES likely tied to Opus 5's 56.3%/64.7% headline figures, but 7-day pullback (-4%) suggests reassessment of official-leaderboard lag/methodology risk.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- claude_news synthesis: resolution likely hinges on protocol (no-tools vs with-tools) and which leaderboard source is used; under with-tools, 60% already met (64.7%), but strict official no-tools leaderboard has not confirmed this.
- code_execution Monte Carlo: raw P(≥60%) estimate ~8-15% under conservative ceiling assumptions purely from official-leaderboard trend extrapolation — much lower than Polymarket's 60%, reflecting divergent read of "official" vs "real-world" Claude capability.
# Directional lean per outcome
- **Yes**: Opus 5 already at 56.3% no-tools (self-reported) — only 3.7pp from threshold; further Anthropic releases (Opus 5.x, Sonnet 6, etc.) plausible before Dec 2026; with-tools score already exceeds 60% (64.7%) creating ambiguity that could favor YES if "HLE Accuracy" on leaderboard includes an agentic/tool-use column reported ≥60%.
- **No**: Official Scale/CAIS leaderboard (the explicit resolution source) still shows sub-40% Claude scores as of Aug 2026, well below Opus 5's self-reported number — indicating either measurement discrepancy or delayed/never-added entries; historical HLE growth has been decelerating/plateauing (Opus 4.1 regression); trend-based statistical models put probability in single-to-low-double digits.
# Gaps / unknowns
- Unclear whether/when Scale AI will add Opus 5 (or later models) to the official leaderboard, and whether its official score will match the 56.3% self-reported figure or be lower (as historical official vs third-party gaps suggest).
- No clarity on whether "HLE Accuracy" per market rules could refer to a tool-augmented score if that becomes the leaderboard's primary displayed metric.
- No confirmed Claude releases/scores for Sept–Dec 2026.
# Calibration anchors
- Polymarket YES price (anchor): 60% (volatile, 17%-64.5% range over 29 days).
- Statistical/trend-based model estimate: ~8-15% (independent of market sentiment).
- Precedent: Opus 4.1 regressed vs Opus 4 on HLE, illustrating non-monotonic single-model progress; official leaderboards have consistently lagged and undercut headline/self-reported figures by ~15-20pp in 2026 snapshots.