← Back to scans

Will the highest score achieved by a SpaceXAI Grok model on Humanity’s Last Exam in 2026 be 45% or higher?

0xeb5d918348e57c731bc6899c18f0476cba775332fcdb8ee53f23a372e777db19 · Science and Technology · 2026-08-21
32%
Agent
78%
Market Price
-46.0%
Edge
48%
Confidence
Volume: 15,224
Spread: 20.0c
Days to resolution: 132
Markets in event: 5
Final Rationale
The event as scoped in the brief requires a Grok model to reach ≥45% HLE Accuracy on the official agi.safe.ai/Scale leaderboard, and that framing sharply favors No: standard Grok sits at ~24.5–27% no-tools, xAI's 44.4%/50.7% Grok 4 Heavy figures (tool-augmented, text-only subset) have never been reproduced or accepted there, and Grok 5 — the only plausible source of a step-change — remained undelivered as of Aug 2026 with flagship Grok 4.6 showing no reported HLE >45%. Even a late-2026 Grok 5 launch would need benchmarking plus leaderboard submission/acceptance before Dec 31, adding lag risk neither forecaster modeled, while the bar itself keeps rising (rivals moved 37.5%→55%). The thin, volatile ~70% Polymarket price ($15k volume, 46.5–97.75% range, falling) reflects mainly the permissive/self-reported interpretation and deserves heavy discounting, especially given the near-0% pricing on the analogous 'Grok 4 above 40%' contract. I therefore side closer to Forecast 2 and the critique, retaining ~30% for a loose resolution or a surprise Grok 5 ship-and-list before year end.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 12$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price and price history for this market, and are there sibling markets at other thresholds (e.g., 35%, 40%, 50%) that imply a distribution?
  2. What score does the official Humanity's Last Exam leaderboard at agi.safe.ai currently list for the best Grok model (Grok 4 / Grok 4.1 / Grok 4.1 Fast), and does it list no-tools or tool-augmented accuracy?
  3. What are the highest reported HLE accuracies for any frontier model as of now (GPT-5.1/5.2, Gemini 3 Pro/Deep Think, Claude Opus 4.5), and how fast has the state of the art on HLE risen month-over-month since Jan 2025?
  4. Has xAI announced or shipped Grok 5, and what HLE performance has xAI claimed or been benchmarked at for its newest models?
  5. How frequently and how quickly does the agi.safe.ai leaderboard add new models (i.e., will a new Grok model actually be scored there in 2026)?
  6. Do xAI's self-reported HLE numbers (e.g., Grok 4 Heavy 44.4%/50.7% with tools) get reproduced on the official leaderboard, or does the leaderboard show notably lower no-tool figures?
Planner reasoning
This is a Polymarket question about whether any Grok model reaches ≥45% HLE accuracy on the official agi.safe.ai leaderboard during 2026. Key drivers: what Grok currently scores on that specific leaderboard, how fast frontier HLE scores are rising (GPT-5.x, Gemini 3 already reported >40-50% with tools), whether xAI ships Grok 5 in 2026, and whether the leaderboard's listed metric is tool-augmented or no-tools. I'll anchor on the Polymarket price, then gather leaderboard/benchmark news and cross-market signals.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will the highest score achieved by a SpaceXAI Grok model on Humanity’s Last Exam in 2026 be 45% or higher?** - Current price (probability): 69.95% - 7-day price change: -8.55% - 30-day price change: +23.45% - Total volume: $15,224 (USD notional) - Price range: 46.
polymarket_related OK 2.1s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Grok': 0 markets | keyword 'HLE': 0 markets | keyword 'xAI': 0 markets
kalshi_related OK 1.9s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'Grok': no matches | keyword 'AI benchmark': ok
claude_news OK 29.7s 11 Here are key findings on the SpaceXAI/xAI Grok HLE benchmark situation: **Official/original scores (July 2025 xAI launch)** - xAI's own launch materials state Grok 4 Heavy saturates most academic benchmarks and is the first model to score 50% on Humanity's Last Exam , though this is specifically
claude_news OK 29.1s 10 **Key findings on HLE trend & xAI/Grok status (as of research cutoff, mid-2026):** - **xAI's self-reported Grok 4 HLE scores (July 2025 launch):** Grok 4 Heavy saturates most academic benchmarks and is the first model to score 50% on Humanity's Last Exam , and Grok 4 Heavy leads USAMO'25 with 61.
gdelt_news OK 239.3s 10 GDELT: 10 articles across 3 queries (lookback=90d). "Grok Humanity's Last Exam benchmark": error GDELT rate-limited after retries (429) | 'Grok 5 xAI release benchmark': 10 hits | "Humanity's Last Exam leaderboard record score": error GDELT rate-limited after retries (429)
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 52.7s 0 ## Key Findings from Quantitative Analysis **Historical anchor points used:** - Grok 3 (~Feb 2025): ~11.5% no-tools - Grok 4 (Jul 2025): 25.4% no-tools / 38.6% with tools - Grok 4 Heavy (Jul 2025): 44.4–50.7% with tools (already spans the 45% threshold) - Grok 4.1 (~Nov 2025): ~27% no-tools / ~46.5
3. Evidence Brief Sonnet · 7171 chars
# Current state The official Humanity's Last Exam leaderboard (Scale AI/CAIS, agi.safe.ai) lists standard Grok 4 at only ~24.5–27% (no-tools); xAI's marketing claim of 44.4%/50.7% (tool-augmented, text-only subset) for "Grok 4 Heavy" has never been reproduced on that official leaderboard. As of the most recent snapshots (Q1–Q3 2026), the leaderboard's 45%+ tier is occupied by Gemini/GPT/Claude models, not Grok. Grok 5 — which Musk suggested could be "nearly perfect" on HLE — has not shipped; xAI's flagship as of Aug 2026 is Grok 4.6. # Timeline of key events - 2025-07 (confirmed): xAI launches Grok 4; self-reports 25.4% (no tools), 38.6% (tools), and Grok 4 Heavy at 44.4%/50.7% (tools, text-only subset). None of the Heavy figures appear on the official leaderboard at launch. - 2025-11 (confirmed): Grok 4.1 released (~27% no-tools per some trackers); Gemini 3 Pro launches, becomes leaderboard leader at 37.5% (no tools). - 2026-01 (rumored): Musk states Grok 5 could be "nearly perfect" on HLE and claims Grok 4 scored ~52% "excluding visual questions" — unverified oral claim, not published/audited. - 2026-02 (confirmed): Gemini 3 Deep Think sets new leaderboard standard at 48.4% (no tools). - 2026-02/03 (confirmed): SpaceX completes acquisition of xAI, forming "SpaceXAI." - 2026-03 (reported): Scale AI leaderboard snapshot: Gemini 3.1 Pro 46.44%, GPT-5.4 Pro 44.32%, Claude Opus 4-7 36.20%; Grok absent from top entries. - 2026-05/06 (reported): Grok 4.5 enters private testing (SpaceX/Tesla engineers involved); parameter count reportedly tripled, coding-focused. - 2026-06 (reported): Grok 5 slips past Q1 2026 target; xAI points to Q2 2026. - 2026-08-12/13 (confirmed): Grok 4.6 launches, replacing Grok 4.5 as flagship; no official HLE score >45% reported. - 2026-08 (reported, third-party aggregators, not the official leaderboard): pricepertoken.com/felloai.com list Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) atop HLE; no Grok model mentioned in the top tier. # Event Will any SpaceXAI Grok model reach ≥45% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct price was returned in research for this ticker. Best available anchor is this market's own Polymarket price: **69.95% YES**, down 8.55% over 7 days but up 23.45% over 30 days; range 46.5%–97.75% over 29 data points; thin volume (~$15.2k total). Treat as noisy/illiquid consensus, not a hardened Kalshi anchor. # Sub-question answers 1. **Polymarket price/siblings** — Current price 69.95% YES (down from 97.75% high); no sibling threshold markets (35/40/50%) found for this specific "any Grok model, 2026" framing. [polymarket_direct]. A separately-worded market "Grok 4 scores above 40% on HLE" (a *different*, likely resolved/narrower question about Grok 4 specifically) shows 0% [claude_news] — not directly comparable. 2. **Official leaderboard Grok score** — Standard Grok 4 sits at ~24.5% (no-tools) on the Scale AI/CAIS leaderboard as of early-2026 snapshots; no Grok model appears in the leaderboard's 44-48% top tier. [intuitionlabs.ai; labs.scale.com] 3. **Frontier SOTA trajectory** — Gemini 3 Pro 37.5% (Nov 2025) → Gemini 3 Deep Think 48.4% (Feb 2026) → Gemini 3.1 Pro 46.44%/GPT-5.4 Pro 44.32% (~Mar 2026) → third-party aggregators show Claude Fable 5 55.5%/GPT-5.6 Sol 49.5% (Aug 2026). SOTA rose ~10-18 points in ~9 months. [blog.google; labs.scale.com; pricepertoken.com; felloai.com] 4. **Grok 5 status** — Not released as of mid/late-2026; repeatedly delayed (Q1→Q2 2026, then further). Flagship is Grok 4.6 (Aug 2026). Musk claimed (unverified) Grok 5 could be "nearly perfect" on HLE. [nxcode.io; latestly.com; Wikipedia] 5. **Leaderboard update cadence** — Frontier models (Gemini, GPT, Claude) get added within weeks of release; Grok 4/4.1 were added but only at their standard (lower) no-tools scores — xAI's Heavy/tool-augmented figures were never submitted/accepted. No evidence a new Grok model has cracked 45%+ on the official board through the research window. 6. **Self-reported vs. official reproduction** — Not reproduced. xAI's 44.4%/50.7% figures (tool-augmented, text-only subset) remain unmatched on the official full-benchmark leaderboard, where Grok sits at 24-27% (no-tools). [Scientific American; intuitionlabs.ai] # Key facts (high-confidence, factual) 1. [x.ai] Grok 4 Heavy self-reported 44.4%/50.7% (tools, text-only subset), July 2025. 2. [Scientific American] These Heavy figures had not appeared on the official leaderboard at launch. 3. [intuitionlabs.ai/labs.scale.com] Official leaderboard shows standard Grok 4 near 24.5-27%, with no Grok model in the ~45%+ top tier through ~March 2026. 4. [blog.google] Gemini 3 Deep Think reached 48.4% (no tools) officially by Feb 2026. 5. [nxcode.io] Grok 5 delayed multiple times; undelivered through mid-2026. 6. [Wikipedia] Grok 4.5 (2026) and Grok 4.6 (Aug 2026) are the newest shipped models; SpaceX absorbed xAI in Feb 2026. # Cross-market signals - Kalshi related: no direct Grok/HLE match found; unrelated markets only. - Polymarket (this market): 69.95% YES, volatile, thin volume, declining last week. - Polymarket (differently-scoped Grok/HLE market): near 0% — but not the same question, limited comparability. # Analyst opinions and speculation - Monte Carlo/code-execution model: probability swings from ~97% (if "official" score accepts tool-augmented figures) to ~20-35% (if leaderboard strictly requires no-tools scores) — a single definitional issue dominates the estimate. Blended estimate offered: 75-90%, but this leans heavily on treating xAI's contested Heavy figure as decisive, which conflicts with the official-leaderboard facts above. - Musk's Jan 2026 claim that Grok could be "nearly perfect" on HLE is promotional/unverified, not evidence of leaderboard reality. # Directional lean per outcome - **Yes**: Only supported if leaderboard rules count tool-augmented scores or Grok 5 delivers a large single-generation jump; Musk's claims and 2025 Heavy figures give some precedent for such framing. - **No**: Official leaderboard has consistently shown Grok well below 45% (24-27% no-tools) even as rivals (Gemini, GPT, Claude) crossed 45-55%; Grok 5 (the model most likely to leap) remains undelivered as of Aug 2026; xAI's high self-reported figures have never been validated on agi.safe.ai. # Gaps / unknowns - No confirmed Kalshi YES price for this exact ticker was retrieved. - Unclear whether agi.safe.ai will ever accept/list tool-augmented scores as the "HLE Accuracy" metric. - No confirmation Grok 4.5/4.6 have been benchmarked on HLE at all, officially or self-reported. - Grok 5 release date and its HLE performance remain unknown/speculative. # Calibration anchors - Polymarket YES price for this exact contract: 69.95% (volatile, thin volume) — best available cross-market anchor absent Kalshi data. - Precedent: prior identically-styled Grok/HLE threshold market ("Grok 4 above 40%") resolved/priced near 0%, reflecting how self-reported tool-augmented Grok scores have failed to translate into official leaderboard confirmations.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.38
Yes 61%
No 39%
The outcome hinges almost entirely on a definitional issue: under a strict reading of the official agi.safe.ai/Scale leaderboard (no-tools, full benchmark), Grok has consistently sat at ~24–27% while Gemini/GPT/Claude crossed 45–55%, and Grok 5 — the only plausible source of a big jump — remained undelivered as of Aug 2026 with the flagship being Grok 4.6 and no reported HLE >45%. Under a looser reading that accepts xAI's self-reported/tool-augmented figures, Yes is near-automatic since Grok 4 Heavy already claimed 44.4%/50.7% back in July 2025 and any 2026 successor would exceed that. The thin, volatile Polymarket anchor at ~70% YES (down from 97.75%, off 8.55% in a week) suggests the market leans toward the permissive interpretation but is losing conviction as 2026 progresses without a qualifying Grok result. Blending ~95% under loose resolution with ~25–30% under strict official-leaderboard resolution lands slightly below the Polymarket price.
gpt-5.5
0.57
Yes 35%
No 65%
Polymarket’s exact-market anchor is 69.95% YES, but it is thin and volatile, so I discount it substantially. The official leaderboard evidence points toward No because Grok’s reproduced no-tools score is only about 24.5–27%, while no Grok model appears in the official 45%+ tier. The frontier trajectory points somewhat toward Yes because Gemini/GPT/Claude models have crossed or approached 45–55%, showing the benchmark is now reachable by leading labs, but Grok 5 remains delayed and Grok 4.6 has no reported official 45%+ score. The self-reported Grok 4 Heavy 44.4%/50.7% figures point only weakly toward Yes because they were tool-augmented/text-subset claims and have not been accepted on agi.safe.ai.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts hedge heavily on a "definitional issue" (loose vs strict resolution), but the event text explicitly specifies "on the official agi.safe.ai leaderboard" — this should sharply constrain weight toward the strict interpretation rather than a near-50/50 blend; Forecast 1's 0.61 in particular seems to overweight the loose/tool-augmented reading despite the brief noting it "has never been reproduced on that official leaderboard." 2. Neither forecaster adequately incorporates the directly comparable precedent cited in the brief: "a prior identically-styled Grok/HLE threshold market ('Grok 4 above 40%') resolved/priced near 0%" — this is strong evidence against the market's own tendency to overprice Grok thresholds, yet both anchor materially above this precedent (0.61 and 0.35). 3. Both treat the Polymarket 69.95% price as a meaningful anchor despite the brief's explicit warning it is "noisy/illiquid consensus" with only ~$15.2k volume and a 46.5%-97.75% range across just 29 data points — Forecast 1 explicitly lands "slightly below" this thin price, which risks anchoring bias rather than independent analysis. 4. Neither forecaster seriously engages with Grok 4.5/4.6 benchmark uncertainty — the brief notes "No confirmation Grok 4.5/4.6 have been benchmarked on HLE at all, officially or self-reported" — meaning even a favorable Grok 5 release still requires a leaderboard submission/acceptance step by Dec 31, 2026, which given "leaderboard update cadence" evidence could itself introduce lag risk not modeled. 5. Both underweight the SOTA trajectory context cutting against Grok even under a "big jump" scenario: rivals have moved from 37.5% to 55.5% in nine months, meaning even if Grok 5 ships and clears 45%, it would be catching up to a rapidly rising bar, not a static one — this dynamic-target risk is mentioned in the brief's "Frontier SOTA trajectory" but not reflected in either probability. 6. Forecast 1's 0.38 confidence and 0.61 Yes and Forecast 2's 0.57 confidence and 0.35 Yes disagree by 26 points with no reconciliation of why one should trust the market anchor more than the hard leaderboard data — this spread itself signals the definitional ambiguity should have been resolved by re-reading the event's explicit "official agi.safe.ai leaderboard" framing rather than left as a wide unresolved split.
6. Aggregator Opus · final calibrated forecast
32% (48% confidence)
Yes: 32% No: 68%
The event as scoped in the brief requires a Grok model to reach ≥45% HLE Accuracy on the official agi.safe.ai/Scale leaderboard, and that framing sharply favors No: standard Grok sits at ~24.5–27% no-tools, xAI's 44.4%/50.7% Grok 4 Heavy figures (tool-augmented, text-only subset) have never been reproduced or accepted there, and Grok 5 — the only plausible source of a step-change — remained undelivered as of Aug 2026 with flagship Grok 4.6 showing no reported HLE >45%. Even a late-2026 Grok 5 launch would need benchmarking plus leaderboard submission/acceptance before Dec 31, adding lag risk neither forecaster modeled, while the bar itself keeps rising (rivals moved 37.5%→55%). The thin, volatile ~70% Polymarket price ($15k volume, 46.5–97.75% range, falling) reflects mainly the permissive/self-reported interpretation and deserves heavy discounting, especially given the near-0% pricing on the analogous 'Grok 4 above 40%' contract. I therefore side closer to Forecast 2 and the critique, retaining ~30% for a loose resolution or a surprise Grok 5 ship-and-list before year end.
Pipeline Timing
Total pipeline time: 365.8s
Per-tool research timings shown in the Research section above.