# Current state
The market resolves on the **official agi.safe.ai / Scale AI HLE leaderboard**, which as of August 2026 shows Anthropic's best Claude model (claude-opus-4-6-thinking-max / Opus 4.6, or Opus 4.7 in combined mode) topping out around **36–40%** — well below 60%. Anthropic's own self-reported model-card and third-party aggregator numbers (Artificial Analysis, llm-stats, BenchLM) claim much higher "with tools" scores for Claude Opus 5 (56.3% no-tools / 64.7% with tools, per Anthropic's July 2026 release), but these are NOT yet reflected on the resolution-source leaderboard, creating a critical protocol/source discrepancy.
# Timeline of key events
- 2025 (throughout): HLE scores near-zero for early models (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%, o1 8.0%) — confirmed, historical baseline [intuitionlabs.ai].
- 2025-11: Gemini 3 scores 37.5% no-tools / 45.8% with tools on HLE — confirmed via multiple aggregators [sentisight.ai].
- 2026-06-10/12: Claude Fable 5 released; reported as "smartest" model but noted media caution about benchmark comparability — confirmed release, reported framing [livemint.com, timesofindia.com].
- 2026-07-24/25: Claude Opus 5 released; Anthropic self-reports 56.3% (no tools) / 64.7% (with tools) on HLE — reported, per Anthropic materials via MarkTechPost, corroborated by llm-stats.com and BenchLM.ai (confirmed as *self-reported/aggregator* figures, not yet on official Scale AI leaderboard).
- 2026-07-09: Commentators flag conflict-of-interest concerns re: Meta Muse Spark's HLE claims — reported, signals broader benchmark-gaming skepticism [digg.com].
- 2026-08 (mid-late): Official Scale AI/agi.safe.ai leaderboard (text-only, no-tools) shows top Claude entry (Opus 4.6-thinking-max) below Gemini 3.1 Pro (47.31%) and GPT-5.4-Pro (45.32%); combined (with-tools) leaderboard shows top Claude (Opus 4.7) at 36.20% — confirmed, primary resolution source.
- Ongoing: Independent verification (Artificial Analysis) of Opus 5 without Anthropic's own protocol puts scores closer to 53% — reported, suggests self-reported 64.7% may not replicate under standardized/independent conditions.
# Event
Will any Anthropic Claude model reach ≥60% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct data was returned; the only direct market pricing available is **Polymarket: 58.5% YES** (down 5.5% over 7 days, up 7.0% over 30 days, range 17–65%, $33.4K volume, 38 data points). Treat this as the crowd anchor in absence of Kalshi data; note it is notably higher than what official-leaderboard evidence supports.
# Sub-question answers
1. **Current highest official Claude HLE score vs. Anthropic's own reporting** — Official Scale AI/agi.safe.ai leaderboard: Claude tops out ~36–40% (Opus 4.6/4.7). Anthropic's Opus 4.6 system card claims "leads all frontier models" at 40.0% no-tools; Opus 5 (July 2026) self-reports 56.3% no-tools/64.7% with tools [claude_news, Anthropic materials].
2. **No-tools vs. tool-augmented leaderboard structure** — agi.safe.ai/Scale AI hosts both a text-only no-tools leaderboard and a combined (with-tools) leaderboard; the gap for frontier models is large (~10-20 points), e.g., Gemini 3: 37.5%→45.8% with tools [claude_news].
3. **Overall SOTA as of early-to-mid 2026** — Gemini 3.1 Pro leads no-tools official leaderboard at 47.31%, GPT-5.4-Pro at 45.32% (Aug 2026); frontier moved from ~37.5% (Nov 2025) to ~47% (Aug 2026) — a slower pace than the 2025 near-zero-to-40% jump [claude_news].
4. **Leaderboard update cadence/risk of Claude being unlisted** — Leaderboard appears actively maintained (multiple Claude and competitor entries current through Aug 2026), so omission risk is low, but there's a lag: newer models (e.g., Opus 5) may not yet appear on official rankings despite public release.
5. **2026 Anthropic release cadence** — Opus 4.6 (mid-2026) → Opus 5 (Jul 2026) → Fable 5/Mythos (mid-2026, restricted). Successive generations reportedly deliver large jumps in self-reported (non-official) HLE scores (30.8%→40.0%→56.3% no-tools across Opus 4.5→4.6→5).
6. **Sibling markets/crowd distribution** — No Polymarket sibling markets found (0 matches for HLE/Claude/Anthropic keywords). Kalshi-related markets are unrelated (IPO race, ERISA, TV release). No cross-market distributional signal available.
# Key facts (high-confidence, factual)
1. [claude_news] Official Scale AI leaderboard: top Claude no-tools ~36-40%, combined/tools ~36.2% (Opus 4.7), as of Aug 2026 — far below 60%.
2. [claude_news/MarkTechPost] Anthropic self-reports Opus 5 at 56.3% (no tools)/64.7% (with tools) — not yet confirmed on official leaderboard.
3. [Artificial Analysis via trendingtopics.eu] Independent re-evaluation of Opus 5 shows ~53%, below Anthropic's self-reported figure.
4. [polymarket_direct] Market price 58.5%, volatile (range 17-65%), suggesting substantial disagreement/uncertainty.
5. [code_execution model] Monte Carlo estimate: P(YES) ≈ 25-36%, highly sensitive to whether resolution counts tool-augmented scores and whether Claude tracks frontier pace.
# Cross-market signals
- Kalshi related: no directly relevant markets found (IPO race, ERISA, unrelated).
- Polymarket: only this market itself; no sibling threshold markets (40%/50%/70%) found to triangulate distribution.
- Sportsbook implied: none applicable.
# Analyst opinions and speculation
- claude_news synthesis flags this as fundamentally a **protocol/definition dispute**: if resolution uses Anthropic's own "with tools" self-reported numbers, threshold already met; if using official Scale AI no-tools/combined leaderboard (the stated resolution source), threshold is far from met (~36-40%).
- code_execution model’s calibrated point estimate (~25-30%) explicitly discounts self-reported Anthropic figures given Claude's documented historical underperformance vs. OpenAI/Google on HLE specifically.
# Directional lean per outcome
- **Yes**: Anthropic's own Opus 5 claims (64.7% with tools) already exceed threshold; rapid 2025-2026 generational jumps (30.8%→56.3% no-tools in ~1 year) suggest further gains plausible; Polymarket prices it fairly likely (58.5%).
- **No**: The actual official resolution source (agi.safe.ai/Scale AI) shows Claude far below 60% (~36-40%) even in tool-augmented mode; independent verification of Opus 5 undercuts Anthropic's self-reported 64.7%; frontier-wide official no-tools leaderboard SOTA is only ~47% (non-Claude); benchmark saturation/deceleration risk.
# Gaps / unknowns
- Whether Scale AI/agi.safe.ai will list Opus 5 or later models with tool-augmented configurations matching Anthropic's internal setup, and whether resolved scores will replicate the 64.7% claim.
- No Kalshi-direct price was returned in research; anchor relies on Polymarket only.
- No visibility into planned Claude 5-series (Opus 6?) roadmap beyond Opus 5 for remainder of 2026.
# Calibration anchors
- Polymarket YES price (used as anchor in absence of Kalshi data): 58.5%.
- Model-based estimate: ~25-36% (code_execution Monte Carlo), reflecting weight on official/no-tools leaderboard evidence.
- Precedent: benchmarks (e.g., MMLU, GPQA) often show 6-12 month lag between vendor self-reported SOTA claims and independent/official leaderboard confirmation.