# Current state
The official/tracked HLE leaderboards (Scale AI labs.scale.com, mirrored on Wikipedia) currently show Gemini 3.1 Pro Preview as Google's best-scoring model at **46.44% accuracy (no-tools)**, still below the 50% threshold. Rival labs (Anthropic's Claude Fable 5/Opus 5, OpenAI's GPT-5.6 Sol) have already crossed 50% per some trackers, proving the bar is reachable, but Gemini has not yet done so with ~4.5 months remaining until the Dec 31, 2026 resolution deadline.
# Timeline of key events
- 2025-01: HLE launches; top model scores ~9% (confirmed, per Wikipedia/code_execution historical data).
- 2025-06: Gemini 2.5 Pro/Deep Think reaches ~21.6% (reported, code_execution synthesis).
- 2025-11: Gemini 3 Pro launches at 37.5% (no tools), beating GPT-5.1's 26.5% (confirmed, blog.google).
- 2026-02: Gemini 3 Deep Think launches at 41.0% (no tools); later reports cite an updated 48.4% (no tools) (reported, Google blog/Remio.ai).
- 2026-02-19: Gemini 3.1 Pro reported at ~44.4–46.44%; described as not beating Claude Opus 4.6 (reported, Substack/Texas A&M).
- 2026-06 to 2026-08: Gemini 3.5 Pro launch delayed (base-model transition, reasoning shortcomings); interim Flash-tier releases (3.5 Flash, 3.6 Flash, 3.7 Flash) ship without notable HLE claims (reported, TechTimes/Geeky Gadgets/Fonearena).
- 2026-08 (recent): Scale AI leaderboard snapshot shows Gemini 3.1 Pro Preview at 46.44% (top Gemini, #1 overall on that board) (confirmed-ish, Scale/Wikipedia). PricePerToken/BenchLM report Claude and OpenAI models above 50–65% (reported, methodology/tool-use unclear).
- 2026-08: Gemini 4 confirmed as a training run only; Pichai states ambition for a "larger base model," no release date; rumors point to Nov/Dec 2026 (reported, emergent.sh/AIToolsReview).
# Event
Will any Google Gemini model reach ≥50% accuracy on the official Humanity's Last Exam leaderboard by Dec 31, 2026?
# Outcomes to forecast
- Yes
- No
# Kalshi market anchor
**No kalshi_direct price was returned in research** — this is a gap. The only direct pricing data available is from Polymarket (same underlying event, likely mirrored/aggregated): **YES ≈ 70.05%**, up +9.05% over 30 days, range 53.05%–96.85%, volume $15,118. Treat this as the best available consensus proxy in place of a missing Kalshi anchor.
# Sub-question answers
1. **Current highest Gemini HLE score / tool inclusion?** Best confirmed Gemini score is 46.44% (Gemini 3.1 Pro Preview, "thinking high," no tools) per Scale AI leaderboard/Wikipedia mirror (claude_news). Resolution source is agi.safe.ai specifically, not directly queried in this research — a gap.
2. **Gemini 3 Pro / Deep Think / later scores?** Gemini 3 Pro: 37.5% (no tools, Nov 2025); Gemini 3 Deep Think: 41.0% initially, later updated to 48.4% (no tools, Feb 2026); Gemini 3.1 Pro: 44.4–46.44% (no tools, varies by tracker). No tool-augmented Gemini scores are cited in sourced material.
3. **HLE growth rate & extrapolation?** ~9% (Jan 2025) → ~46% (Aug 2026), averaging ~2–3 pts/month early on, decelerating recently. Code-execution Monte Carlo modeling gives wide range (55–85%+ under various fits); a discounted judgment estimate lands at ~60–70% probability of Gemini reaching ≥50% by Dec 2026.
4. **Has any model exceeded 50%?** Yes — Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) per PricePerToken (Aug 2026); BenchLM shows even higher (Opus 5 at 64.7%), likely reflecting tool-use/methodology differences (unreconciled discrepancy).
5. **2026 Gemini releases/benchmarks previewed?** Gemini 3.5 Pro delayed (base-model rebuild, reasoning issues); interim Flash-tier models (3.5/3.6/3.7 Flash) shipped without major HLE claims; Gemini 4 remains in training only as of Aug 2026, with no confirmed release date (rumored Nov/Dec 2026).
6. **Leaderboard update frequency/lag?** Not explicitly documented; one snapshot was cached "~36 minutes" prior to query, suggesting frequent updates. Multiple third-party trackers (Scale, PricePerToken, BenchLM, Artificial Analysis) show inconsistent numbers for the same models — meaningful resolution-source risk.
7. **Polymarket price for this and related markets?** This exact market: 70.05% YES on Polymarket. A separate "Gemini 3 Predictions" page cites a "50%+" outcome near 73% (claude_news, unverified page). No 60%/70% threshold sibling markets were found via polymarket_related scan (0 matches).
# Key facts (high-confidence, factual)
1. [Google blog] Gemini 3 Pro scored 37.5% (no tools) at Nov 2025 launch.
2. [Google blog/Remio.ai] Gemini 3 Deep Think scored 41.0%, later reported at 48.4% (no tools, Feb 2026).
3. [Scale AI/Wikipedia] Gemini 3.1 Pro Preview currently tops Gemini scores at 46.44% (no tools), as of ~Aug 2026.
4. [PricePerToken] Claude Fable 5 (55.5%) and GPT-5.6 Sol (49.5%) have already reached/exceeded 50% on at least one tracker.
5. [emergent.sh] Gemini 4 is in training only as of Aug 2026; no release date confirmed.
# Cross-market signals
- Kalshi related: no direct Gemini/HLE match found; unrelated markets only.
- Polymarket: 70.05% YES on this exact market, uptrending (+9pp/30d); a related "Gemini 3 Predictions" page cites ~73% for a "50%+" outcome (unverified, likely same/related market).
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Neurosciencenews frames 50% as a real ceiling: "Even the most advanced current models struggle to exceed 50%."
- Code-execution quantitative model, after discounting logistic overfit, lands at ~60–70% probability of crossing 50%, with tool-augmented scoring inclusion being the key swing factor.
- FutureHouse contamination critique (~30% of chem/bio answers possibly wrong) could distort future scoring dynamics either direction.
# Directional lean per outcome
- **Yes**: Steady multi-model upward trend (37.5%→41%→48.4%→46.44%); rivals already past 50%, proving feasibility; Gemini 4 possible before year-end; Polymarket pricing at 70%.
- **No**: Current best Gemini score (46.44%) still short by ~3.5+ points; Gemini 3.5 Pro delayed with quality issues; Gemini 4 unconfirmed/training-only with no near-term date; recent trajectory shows deceleration/non-monotonicity (46.44% is actually below Deep Think's reported 48.4%).
# Gaps / unknowns
- No kalshi_direct price was retrieved — primary anchor is missing; relying on Polymarket proxy.
- Direct agi.safe.ai leaderboard values not fetched; reliance on Scale AI/Wikipedia mirrors introduces sourcing uncertainty.
- Discrepancy between "48.4%" Deep Think claim and "46.44%" current leaderboard top is unreconciled (possibly different scoring runs/methodologies).
- Tool-augmented score inclusion on the resolution leaderboard is unconfirmed — critical to the 50% threshold.
# Calibration anchors
- Polymarket YES price (proxy for missing Kalshi anchor): 70.05%, uptrending.
- Precedent: benchmark scores in this AI cycle have shown fast, non-monotonic jumps (9%→46% in ~19 months), but also stalls/plateaus per-model; closing a 3.5pt gap within ~4.5 months is plausible given historical pace.