← Back to scans

Will the highest score achieved by a SpaceXAI Grok model on Humanity’s Last Exam in 2026 be 45% or higher?

0xeb5d918348e57c731bc6899c18f0476cba775332fcdb8ee53f23a372e777db19 · Science and Technology · 2026-08-28
70%
Agent
74%
Market Price
-4.4%
Edge
48%
Confidence
Volume: 21,997
Spread: 2.5c
Days to resolution: 125
Markets in event: 5
Final Rationale
(aggregator returned no rationale)
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 5$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 32% 78% 48%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related claude_news claude_news gdelt_news kalshi_related wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What Grok models currently appear on the official Humanity's Last Exam leaderboard at agi.safe.ai, and what are their listed HLE Accuracy values?
  2. Does the agi.safe.ai leaderboard report no-tools text-only accuracy or tools/agentic accuracy, and how much lower are Grok's no-tools scores (e.g., Grok 4 ~25%) than its with-tools scores (~41-51%)?
  3. What are the current highest HLE accuracies of any frontier model (GPT-5.x, Gemini 3, Claude Opus 4.5) on that same leaderboard, and what is the trend rate of improvement per year?
  4. Has xAI announced or released Grok 5, and what HLE score has been claimed/projected for it?
  5. How quickly does the agi.safe.ai leaderboard get updated with new model releases (i.e., risk that a qualifying Grok model exists but isn't listed by Dec 31, 2026)?
  6. What is the current Polymarket price for this market and for the sibling threshold markets (e.g., 30%, 35%, 40%, 50%) implying the distribution of expected Grok HLE scores?
Planner reasoning
This is a Polymarket question about whether any xAI Grok model reaches ≥45% HLE accuracy on the official agi.safe.ai leaderboard during 2026, so the market price is the primary anchor and the key empirical facts are (a) current Grok scores as listed on that leaderboard, (b) the leaderboard's methodology (no-tools text-only vs. tools-augmented), and (c) Grok 5 release timing and expected capability jump. Frontier HLE scores have risen fast (GPT-5.x/Gemini 3 in the 30-45% range with tools), so the crux is whether the resolution source lists tool-augmented scores and whether xAI ships a Grok 5 that is benchmarked there.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a SpaceXAI Grok model on Humanity’s Last Exam in 2026 be 45% or higher?** - Current price (probability): 73.95% - 7-day price change: +2.90% - 30-day price change: -7.55% - Total volume: $21,997 (USD notional) - Price range: 46.5
polymarket_related OK 2.8s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Grok': 0 markets | keyword 'HLE': 0 markets | keyword 'xAI': 0 markets | keyword 'best AI model': 0 markets
claude_news OK 31.0s 13 **HLE Leaderboard & Grok Findings** - xAI's own July 2025 launch claims stated "Grok 4 Heavy saturates most academic benchmarks and is the first model to score 50% on Humanity's Last Exam." Specifically, Grok 4 Heavy is "the first to score 50.7% on Humanity's Last Exam (text-only subset)." This
claude_news OK 39.2s 11 Key findings on Grok/xAI HLE performance and Grok 5 timeline: - **Grok 4 Heavy already exceeded 45% on HLE in July 2025** (before the 2026 window even began): xAI stated "Grok 4 Heavy saturates most academic benchmarks and is the first model to score 50% on Humanity's Last Exam." Specifically, G
gdelt_news OK 2.9s 0 GDELT: 0 articles across 3 queries (lookback=90d). "Grok Humanity's Last Exam score": error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28grok+OR+humanity+OR+last+OR+exam+OR+score%29+sourcelang%3Aenglish&mode=ArtList&format=json&s
kalshi_related OK 2.6s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Grok': no matches
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 53.2s 0 ## Key Findings - **Historical Grok-specific trajectory (used for fit):** Grok‑3 (~Feb 2025) ≈10% no‑tools → Grok‑4 (Jul 2025) ≈25% no‑tools → Grok‑4 Heavy (Jul 2025) ≈38‑39% no‑tools / ≈44% with tools (search+code). This mirrors the broader lab pattern of +15‑22 pts uplift from tool augmentation.
3. Evidence Brief Sonnet · 7004 chars
# Current state Resolution hinges on the official agi.safe.ai HLE leaderboard reporting a SpaceXAI Grok model's "HLE Accuracy" ≥45% at any point before 2026-12-31. xAI/SpaceXAI has self-reported Grok 4 Heavy scores as high as 50.7% (text-only, agentic) back in mid-2025, but independent trackers and the official leaderboard have historically lagged or diverged from those self-reported figures, and some 2026 leaderboard snapshots show no Grok model in the top bracket at all. # Timeline of key events - 2025-03: Gemini 2.5 Pro launches at ~18.8% HLE (confirmed, historical baseline) — [claude_news]. - 2025-07: xAI launches Grok 4 / Grok 4 Heavy; self-reports Grok 4 no-tools 25.4%, with-tools 38.6%, Grok 4 Heavy with-tools 44.4% (multimodal) and 50.7% (text-only) — confirmed as xAI's own claim [x.ai/news/grok-4, Scientific American]. - 2025-07-17: Manifold market on "Grok 4 lists at 45%+ on official HLE leaderboard" resolves **NO** — confirmed, showing official leaderboard did not corroborate xAI's self-reported number within a month [claude_news]. - 2026-01: Elon Musk states xAI is training Grok 5, targeting 2026 release, claims potential "near-perfect" HLE performance — reported [latestly.com]. - 2026-02: SpaceX acquires xAI (~$250B valuation) — confirmed [Wikipedia]. - Q1 2026: Grok 5 launch window passes without release — reported. - ~2026-06/07: xAI rebrands as SpaceXAI; Grok 5 beta pushed to May–June 2026, full API to Q3 2026 — reported [nxcode.io, overchat.ai]. - 2026 (mid-year, undated snapshot): Grok 4.5/4.6 released; leaderboard snapshots vary — one shows Gemini 3.1 Pro Preview 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Gemini 3 Pro Preview 37.52%, with no Grok in top group [pricepertoken/llm-stats]; Wikipedia separately states Grok 4.6 "scored behind only" Fable 5, Opus 5, and GPT-5.6 Sol — conflicting rank data, reconciled as: Grok's exact standing is uncertain but plausibly near/at the 45% threshold. - 2026-08-27: Leaderboard tops out at Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) — reported; Grok models absent from this top-3 mention [claude_news]. - No confirmed Grok 5 HLE score as of research date; only unverified "Project Valis" leak (45.1% on a different, non-HLE "Zeitgeist" exam) — rumored, low confidence. # Event Will any SpaceXAI Grok model reach ≥45% HLE Accuracy on the official agi.safe.ai leaderboard by 2026-12-31? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct data was returned in this research pull (ticker format matches a Polymarket condition ID, not a native Kalshi ticker). **Primary cross-market anchor is Polymarket direct data for this identical question: current YES price 73.95%**, up +2.90% over 7 days but down -7.55% over 30 days; historical range 46.5%–97.75% over 36 days; volume ~$22K (thin). No Kalshi-specific pricing available — flagged as a gap. # Sub-question answers 1. **Grok models/values on leaderboard** — No single authoritative agi.safe.ai snapshot was retrieved directly; secondary trackers give conflicting pictures: one 2026 snapshot excludes Grok from the 37–46% top band; Wikipedia states Grok 4.6 ranks 4th overall (behind Fable 5, Opus 5, GPT-5.6 Sol at 49.5%+), implying a score plausibly near/above 45% [Wikipedia; pricepertoken/llm-stats]. 2. **No-tools vs tools scoring gap** — Confirmed large gap: Grok 4 no-tools 25.4%, with-tools 38.6%, Grok 4 Heavy with-tools 44.4% (multimodal) / 50.7% (text-only agentic) [Scientific American, x.ai]. 3. **Frontier trend** — Top scores rose from ~18.8% (Mar 2025) to mid-30s (early 2026) to mid-40s (mid-2026) to 49.5–55.5% (Aug 2026, Claude Fable 5/Opus 5, GPT-5.6 Sol) — roughly +30-35 points in 18 months [claude_news]. 4. **Grok 5 status** — Announced/training confirmed (Musk, Jan 2026); release repeatedly delayed (Q1→Q2→beta May/June→API Q3 2026); no confirmed HLE score, only Musk's aspirational "near-perfect" claim and an unverified leak (45.1% on a non-HLE exam) [latestly.com, nxcode.io]. 5. **Leaderboard update lag risk** — Historically significant: Grok 4 Heavy's 50.7% self-reported score did not appear on the official leaderboard within a month (Manifold resolved NO), suggesting real risk that a late-2026 Grok 5 release could miss official listing by Dec 31, 2026 [claude_news]. 6. **Polymarket pricing** — This market: 73.95% YES (polymarket_direct). No sibling threshold markets found (polymarket_related returned 0 matches); a separate "Grok 4 above 40%" Polymarket market was referenced but no price captured. # Key facts (high-confidence, factual) 1. [x.ai/Scientific American] Grok 4 Heavy self-reported 50.7% (text-only) / 44.4% (multimodal, with tools) HLE in July 2025. 2. [claude_news/Manifold] Official leaderboard did not corroborate Grok 4's 45%+ claim within a month of release (Manifold resolved NO, Jul 2025). 3. [Wikipedia] SpaceX acquired xAI Feb 2026; rebranded SpaceXAI; Grok 4.6 described as trailing only the top 3 (Fable 5, Opus 5, GPT-5.6 Sol) as of 2026. 4. [claude_news] As of Aug 2026, top HLE scorers are Claude Fable 5 (55.5%), Opus 5 (54.9%), GPT-5.6 Sol (49.5%); no Grok explicitly in that top-3 mention. 5. [latestly.com/nxcode.io] Grok 5 training confirmed Jan 2026; release date slipped multiple times through 2026, no confirmed HLE score yet. # Cross-market signals - Kalshi related: no direct Grok/HLE matches found. - Polymarket (this exact market): 73.95% YES, volatile (46.5–97.75% range), thin volume (~$22K), 30-day downtrend but 7-day uptick. - One claude_news snippet claims Polymarket shows ~0% YES — contradicts polymarket_direct (73.95%); likely stale/confused with the older Manifold market; direct tool data is treated as authoritative per reconciliation rules. # Analyst opinions and speculation - Monte Carlo/logistic-fit model (code_execution) estimates central probability ~0.45–0.55 after discounting for overfitting and Grok's historical lag behind frontier labs; raw optimistic fit reached ~0.68. - Musk's "near-perfect" Grok 5 claims are promotional and unverified. # Directional lean per outcome - **Yes**: Grok already self-reported >45% in 2025; Grok 4.6 reportedly near top-4; Grok 5 coming with high ambitions; Polymarket prices 74%. - **No**: Official leaderboard has historically lagged/failed to confirm xAI's self-reported numbers; some 2026 leaderboard snapshots exclude Grok from top tier; Grok 5 delays risk missing Dec 31 cutoff; resolution may default to lower no-tools scores (historically ~25-39% for Grok). # Gaps / unknowns - No direct current agi.safe.ai leaderboard reading for a Grok model's exact listed accuracy. - Whether "HLE Accuracy" metric used for resolution is no-tools or tools-inclusive is unresolved and highly consequential. - No confirmed Kalshi-native price exists in this pull. # Calibration anchors - Polymarket YES (proxy anchor): 73.95%. - Precedent: Manifold market on nearly identical question resolved NO in 2025 despite xAI's own 50%+ claim, showing high resolution/reporting risk.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.42
Yes 75%
No 25%
The Polymarket direct price on this identical question (73.95% YES, with a 7-day uptick) is the best available consensus anchor absent Kalshi pricing. Substantive evidence leans YES: Wikipedia describes Grok 4.6 as trailing only Fable 5, Opus 5, and GPT-5.6 Sol (49.5%), which implies a score plausibly in the mid-to-high 40s, and the frontier has moved from ~19% (Mar 2025) to 49.5–55.5% (Aug 2026), so a Grok 4.6/5 listing at ≥45% is well within trend. Counterweights are real but secondary: the July 2025 Manifold precedent showed the official agi.safe.ai leaderboard failing to corroborate xAI's self-reported 50.7%, one 2026 tracker snapshot omits Grok from the 37–46% band, Grok 5's release has slipped repeatedly, and resolution ambiguity between no-tools (historically ~25–39% for Grok) and tools-inclusive scoring could drag the listed figure below 45%. I settle slightly above the Polymarket price's midpoint of its recent volatile range, at 0.75.
gpt-5.6-sol
0.55
Yes 64%
No 36%
The best available market anchor is Polymarket’s 73.95% YES price, though its thin volume, volatility, and lack of a Kalshi-native price reduce its authority. Grok 4 Heavy already self-reported up to 50.7%, and Grok 4.6 is reportedly fourth-ranked, but conflicting snapshots and the large tools-versus-no-tools gap leave its official qualifying score uncertain. Rapid frontier-score improvement and a possible Grok 5 release support YES, while repeated Grok 5 delays and the official leaderboard’s prior failure to corroborate Grok’s 45%+ claims create substantial cutoff and reporting risk. I therefore discount the market consensus to 64% YES.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters anchor heavily on Polymarket's 73.95% but understate the direct evidence conflict noted in the brief itself: one 2026 snapshot explicitly excludes Grok from the top band (Gemini 3.1 Pro 46.44%, GPT-5.4 Pro 44.32%, no Grok listed) while the Aug 2026 leaderboard top-3 also omits Grok entirely (Fable 5, Opus 5, GPT-5.6 Sol only) — this is two independent later-dated snapshots both showing Grok absent from ≥45% territory, which is stronger and more recent evidence than the single Wikipedia "4th place" claim, yet neither forecaster weights this asymmetry explicitly. 2. Neither forecaster adequately addresses the critical unresolved metric ambiguity flagged in the brief: "Whether 'HLE Accuracy' metric used for resolution is no-tools or tools-inclusive is unresolved and highly consequential" — given Grok's historical no-tools scores (~25-39%) are far below 45%, this ambiguity alone could be the single largest swing factor, yet both forecasts treat it as a minor caveat rather than a structural reason to discount below the market price. 3. Both forecasters underweight the Manifold precedent as a repeatable pattern rather than a one-off: the official leaderboard failed to corroborate xAI's self-reported 50.7% for over a month in 2025, and the brief notes "real risk that a late-2026 Grok 5 release could miss official listing by Dec 31, 2026" — if Grok 5 slips to Q3 2026 API (as reported) and needs subsequent leaderboard evaluation/publication lag, a Dec 31 cutoff miss is plausible and deserves more explicit probability mass in the "No" case. 4. Forecast 1's reasoning that Grok 4.6 "trailing only the top 3" implies "a score plausibly in the mid-to-high 40s" is an inference stacked on an inference (Wikipedia's ambiguous ranking language), while the more concrete numeric snapshot (pricepertoken/llm-stats) explicitly shows no Grok in the 37-46% band — Forecast 1 should have flagged this as a direct contradiction requiring resolution rather than implicitly favoring the more optimistic source. 5. Neither forecast explores the tail scenario where SpaceX's acquisition of xAI (Feb 2026) and subsequent "SpaceXAI" rebrand introduces organizational disruption (leadership/engineering focus shift, possible deprioritization of pure LLM benchmarking in favor of SpaceX-integration work) that could further delay Grok 5's release or official benchmarking — a scenario that could push probability toward No but isn't discussed by either forecaster. 6. Both forecasts note Polymarket's thin volume (~$22K) and 30-day downtrend (-7.55%) but don't meaningfully adjust for this — a thinly-traded, volatile market swinging 46.5%-97.75% over just 36 days is weak evidence of true consensus, and the recent trend (down over 30 days, only a small 7-day uptick) arguably supports the case for anchoring lower than 73.95%, which only Forecast 2 partially captures by discounting to 64%.
6. Aggregator Opus · final calibrated forecast
70% (48% confidence)
Yes: 70% No: 30%
(aggregator returned no rationale)
Pipeline Timing
Total pipeline time: 187.3s
Per-tool research timings shown in the Research section above.