# Current state
Resolution hinges on the official agi.safe.ai HLE leaderboard reporting a SpaceXAI Grok model's "HLE Accuracy" ≥45% at any point before 2026-12-31. xAI/SpaceXAI has self-reported Grok 4 Heavy scores as high as 50.7% (text-only, agentic) back in mid-2025, but independent trackers and the official leaderboard have historically lagged or diverged from those self-reported figures, and some 2026 leaderboard snapshots show no Grok model in the top bracket at all.
# Timeline of key events
- 2025-03: Gemini 2.5 Pro launches at ~18.8% HLE (confirmed, historical baseline) — [claude_news].
- 2025-07: xAI launches Grok 4 / Grok 4 Heavy; self-reports Grok 4 no-tools 25.4%, with-tools 38.6%, Grok 4 Heavy with-tools 44.4% (multimodal) and 50.7% (text-only) — confirmed as xAI's own claim [x.ai/news/grok-4, Scientific American].
- 2025-07-17: Manifold market on "Grok 4 lists at 45%+ on official HLE leaderboard" resolves **NO** — confirmed, showing official leaderboard did not corroborate xAI's self-reported number within a month [claude_news].
- 2026-01: Elon Musk states xAI is training Grok 5, targeting 2026 release, claims potential "near-perfect" HLE performance — reported [latestly.com].
- 2026-02: SpaceX acquires xAI (~$250B valuation) — confirmed [Wikipedia].
- Q1 2026: Grok 5 launch window passes without release — reported.
- ~2026-06/07: xAI rebrands as SpaceXAI; Grok 5 beta pushed to May–June 2026, full API to Q3 2026 — reported [nxcode.io, overchat.ai].
- 2026 (mid-year, undated snapshot): Grok 4.5/4.6 released; leaderboard snapshots vary — one shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Gemini 3 Pro Preview 37.52%, with no Grok in top group [pricepertoken/llm-stats]; Wikipedia separately states Grok 4.6 "scored behind only" Fable 5, Opus 5, and GPT-5.6 Sol — conflicting rank data, reconciled as: Grok's exact standing is uncertain but plausibly near/at the 45% threshold.
- 2026-08-27: Leaderboard tops out at Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) — reported; Grok models absent from this top-3 mention [claude_news].
- No confirmed Grok 5 HLE score as of research date; only unverified "Project Valis" leak (45.1% on a different, non-HLE "Zeitgeist" exam) — rumored, low confidence.
# Event
Will any SpaceXAI Grok model reach ≥45% HLE Accuracy on the official agi.safe.ai leaderboard by 2026-12-31?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct data was returned in this research pull (ticker format matches a Polymarket condition ID, not a native Kalshi ticker). **Primary cross-market anchor is Polymarket direct data for this identical question: current YES price 73.95%**, up +2.90% over 7 days but down -7.55% over 30 days; historical range 46.5%–97.75% over 36 days; volume ~$22K (thin). No Kalshi-specific pricing available — flagged as a gap.
# Sub-question answers
1. **Grok models/values on leaderboard** — No single authoritative agi.safe.ai snapshot was retrieved directly; secondary trackers give conflicting pictures: one 2026 snapshot excludes Grok from the 37–46% top band; Wikipedia states Grok 4.6 ranks 4th overall (behind Fable 5, Opus 5, GPT-5.6 Sol at 49.5%+), implying a score plausibly near/above 45% [Wikipedia; pricepertoken/llm-stats].
2. **No-tools vs tools scoring gap** — Confirmed large gap: Grok 4 no-tools 25.4%, with-tools 38.6%, Grok 4 Heavy with-tools 44.4% (multimodal) / 50.7% (text-only agentic) [Scientific American, x.ai].
3. **Frontier trend** — Top scores rose from ~18.8% (Mar 2025) to mid-30s (early 2026) to mid-40s (mid-2026) to 49.5–55.5% (Aug 2026, Claude Fable 5/Opus 5, GPT-5.6 Sol) — roughly +30-35 points in 18 months [claude_news].
4. **Grok 5 status** — Announced/training confirmed (Musk, Jan 2026); release repeatedly delayed (Q1→Q2→beta May/June→API Q3 2026); no confirmed HLE score, only Musk's aspirational "near-perfect" claim and an unverified leak (45.1% on a non-HLE exam) [latestly.com, nxcode.io].
5. **Leaderboard update lag risk** — Historically significant: Grok 4 Heavy's 50.7% self-reported score did not appear on the official leaderboard within a month (Manifold resolved NO), suggesting real risk that a late-2026 Grok 5 release could miss official listing by Dec 31, 2026 [claude_news].
6. **Polymarket pricing** — This market: 73.95% YES (polymarket_direct). No sibling threshold markets found (polymarket_related returned 0 matches); a separate "Grok 4 above 40%" Polymarket market was referenced but no price captured.
# Key facts (high-confidence, factual)
1. [x.ai/Scientific American] Grok 4 Heavy self-reported 50.7% (text-only) / 44.4% (multimodal, with tools) HLE in July 2025.
2. [claude_news/Manifold] Official leaderboard did not corroborate Grok 4's 45%+ claim within a month of release (Manifold resolved NO, Jul 2025).
3. [Wikipedia] SpaceX acquired xAI Feb 2026; rebranded SpaceXAI; Grok 4.6 described as trailing only the top 3 (Fable 5, Opus 5, GPT-5.6 Sol) as of 2026.
4. [claude_news] As of Aug 2026, top HLE scorers are Claude Fable 5 (55.5%), Opus 5 (54.9%), GPT-5.6 Sol (49.5%); no Grok explicitly in that top-3 mention.
5. [latestly.com/nxcode.io] Grok 5 training confirmed Jan 2026; release date slipped multiple times through 2026, no confirmed HLE score yet.
# Cross-market signals
- Kalshi related: no direct Grok/HLE matches found.
- Polymarket (this exact market): 73.95% YES, volatile (46.5–97.75% range), thin volume (~$22K), 30-day downtrend but 7-day uptick.
- One claude_news snippet claims Polymarket shows ~0% YES — contradicts polymarket_direct (73.95%); likely stale/confused with the older Manifold market; direct tool data is treated as authoritative per reconciliation rules.
# Analyst opinions and speculation
- Monte Carlo/logistic-fit model (code_execution) estimates central probability ~0.45–0.55 after discounting for overfitting and Grok's historical lag behind frontier labs; raw optimistic fit reached ~0.68.
- Musk's "near-perfect" Grok 5 claims are promotional and unverified.
# Directional lean per outcome
- **Yes**: Grok already self-reported >45% in 2025; Grok 4.6 reportedly near top-4; Grok 5 coming with high ambitions; Polymarket prices 74%.
- **No**: Official leaderboard has historically lagged/failed to confirm xAI's self-reported numbers; some 2026 leaderboard snapshots exclude Grok from top tier; Grok 5 delays risk missing Dec 31 cutoff; resolution may default to lower no-tools scores (historically ~25-39% for Grok).
# Gaps / unknowns
- No direct current agi.safe.ai leaderboard reading for a Grok model's exact listed accuracy.
- Whether "HLE Accuracy" metric used for resolution is no-tools or tools-inclusive is unresolved and highly consequential.
- No confirmed Kalshi-native price exists in this pull.
# Calibration anchors
- Polymarket YES (proxy anchor): 73.95%.
- Precedent: Manifold market on nearly identical question resolved NO in 2025 despite xAI's own 50%+ claim, showing high resolution/reporting risk.