# Current state
The official Humanity's Last Exam leaderboard (Scale AI/CAIS, agi.safe.ai) lists standard Grok 4 at only ~24.5–27% (no-tools); xAI's marketing claim of 44.4%/50.7% (tool-augmented, text-only subset) for "Grok 4 Heavy" has never been reproduced on that official leaderboard. As of the most recent snapshots (Q1–Q3 2026), the leaderboard's 45%+ tier is occupied by Gemini/GPT/Claude models, not Grok. Grok 5 — which Musk suggested could be "nearly perfect" on HLE — has not shipped; xAI's flagship as of Aug 2026 is Grok 4.6.
# Timeline of key events
- 2025-07 (confirmed): xAI launches Grok 4; self-reports 25.4% (no tools), 38.6% (tools), and Grok 4 Heavy at 44.4%/50.7% (tools, text-only subset). None of the Heavy figures appear on the official leaderboard at launch.
- 2025-11 (confirmed): Grok 4.1 released (~27% no-tools per some trackers); Gemini 3 Pro launches, becomes leaderboard leader at 37.5% (no tools).
- 2026-01 (rumored): Musk states Grok 5 could be "nearly perfect" on HLE and claims Grok 4 scored ~52% "excluding visual questions" — unverified oral claim, not published/audited.
- 2026-02 (confirmed): Gemini 3 Deep Think sets new leaderboard standard at 48.4% (no tools).
- 2026-02/03 (confirmed): SpaceX completes acquisition of xAI, forming "SpaceXAI."
- 2026-03 (reported): Scale AI leaderboard snapshot: Gemini 3.1 Pro 46.44%, GPT-5.4 Pro 44.32%, Claude Opus 4-7 36.20%; Grok absent from top entries.
- 2026-05/06 (reported): Grok 4.5 enters private testing (SpaceX/Tesla engineers involved); parameter count reportedly tripled, coding-focused.
- 2026-06 (reported): Grok 5 slips past Q1 2026 target; xAI points to Q2 2026.
- 2026-08-12/13 (confirmed): Grok 4.6 launches, replacing Grok 4.5 as flagship; no official HLE score >45% reported.
- 2026-08 (reported, third-party aggregators, not the official leaderboard): pricepertoken.com/felloai.com list Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) atop HLE; no Grok model mentioned in the top tier.
# Event
Will any SpaceXAI Grok model reach ≥45% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct price was returned in research for this ticker. Best available anchor is this market's own Polymarket price: **69.95% YES**, down 8.55% over 7 days but up 23.45% over 30 days; range 46.5%–97.75% over 29 data points; thin volume (~$15.2k total). Treat as noisy/illiquid consensus, not a hardened Kalshi anchor.
# Sub-question answers
1. **Polymarket price/siblings** — Current price 69.95% YES (down from 97.75% high); no sibling threshold markets (35/40/50%) found for this specific "any Grok model, 2026" framing. [polymarket_direct]. A separately-worded market "Grok 4 scores above 40% on HLE" (a *different*, likely resolved/narrower question about Grok 4 specifically) shows 0% [claude_news] — not directly comparable.
2. **Official leaderboard Grok score** — Standard Grok 4 sits at ~24.5% (no-tools) on the Scale AI/CAIS leaderboard as of early-2026 snapshots; no Grok model appears in the leaderboard's 44-48% top tier. [intuitionlabs.ai; labs.scale.com]
3. **Frontier SOTA trajectory** — Gemini 3 Pro 37.5% (Nov 2025) → Gemini 3 Deep Think 48.4% (Feb 2026) → Gemini 3.1 Pro 46.44%/GPT-5.4 Pro 44.32% (~Mar 2026) → third-party aggregators show Claude Fable 5 55.5%/GPT-5.6 Sol 49.5% (Aug 2026). SOTA rose ~10-18 points in ~9 months. [blog.google; labs.scale.com; pricepertoken.com; felloai.com]
4. **Grok 5 status** — Not released as of mid/late-2026; repeatedly delayed (Q1→Q2 2026, then further). Flagship is Grok 4.6 (Aug 2026). Musk claimed (unverified) Grok 5 could be "nearly perfect" on HLE. [nxcode.io; latestly.com; Wikipedia]
5. **Leaderboard update cadence** — Frontier models (Gemini, GPT, Claude) get added within weeks of release; Grok 4/4.1 were added but only at their standard (lower) no-tools scores — xAI's Heavy/tool-augmented figures were never submitted/accepted. No evidence a new Grok model has cracked 45%+ on the official board through the research window.
6. **Self-reported vs. official reproduction** — Not reproduced. xAI's 44.4%/50.7% figures (tool-augmented, text-only subset) remain unmatched on the official full-benchmark leaderboard, where Grok sits at 24-27% (no-tools). [Scientific American; intuitionlabs.ai]
# Key facts (high-confidence, factual)
1. [x.ai] Grok 4 Heavy self-reported 44.4%/50.7% (tools, text-only subset), July 2025.
2. [Scientific American] These Heavy figures had not appeared on the official leaderboard at launch.
3. [intuitionlabs.ai/labs.scale.com] Official leaderboard shows standard Grok 4 near 24.5-27%, with no Grok model in the ~45%+ top tier through ~March 2026.
4. [blog.google] Gemini 3 Deep Think reached 48.4% (no tools) officially by Feb 2026.
5. [nxcode.io] Grok 5 delayed multiple times; undelivered through mid-2026.
6. [Wikipedia] Grok 4.5 (2026) and Grok 4.6 (Aug 2026) are the newest shipped models; SpaceX absorbed xAI in Feb 2026.
# Cross-market signals
- Kalshi related: no direct Grok/HLE match found; unrelated markets only.
- Polymarket (this market): 69.95% YES, volatile, thin volume, declining last week.
- Polymarket (differently-scoped Grok/HLE market): near 0% — but not the same question, limited comparability.
# Analyst opinions and speculation
- Monte Carlo/code-execution model: probability swings from ~97% (if "official" score accepts tool-augmented figures) to ~20-35% (if leaderboard strictly requires no-tools scores) — a single definitional issue dominates the estimate. Blended estimate offered: 75-90%, but this leans heavily on treating xAI's contested Heavy figure as decisive, which conflicts with the official-leaderboard facts above.
- Musk's Jan 2026 claim that Grok could be "nearly perfect" on HLE is promotional/unverified, not evidence of leaderboard reality.
# Directional lean per outcome
- **Yes**: Only supported if leaderboard rules count tool-augmented scores or Grok 5 delivers a large single-generation jump; Musk's claims and 2025 Heavy figures give some precedent for such framing.
- **No**: Official leaderboard has consistently shown Grok well below 45% (24-27% no-tools) even as rivals (Gemini, GPT, Claude) crossed 45-55%; Grok 5 (the model most likely to leap) remains undelivered as of Aug 2026; xAI's high self-reported figures have never been validated on agi.safe.ai.
# Gaps / unknowns
- No confirmed Kalshi YES price for this exact ticker was retrieved.
- Unclear whether agi.safe.ai will ever accept/list tool-augmented scores as the "HLE Accuracy" metric.
- No confirmation Grok 4.5/4.6 have been benchmarked on HLE at all, officially or self-reported.
- Grok 5 release date and its HLE performance remain unknown/speculative.
# Calibration anchors
- Polymarket YES price for this exact contract: 69.95% (volatile, thin volume) — best available cross-market anchor absent Kalshi data.
- Precedent: prior identically-styled Grok/HLE threshold market ("Grok 4 above 40%") resolved/priced near 0%, reflecting how self-reported tool-augmented Grok scores have failed to translate into official leaderboard confirmations.