← Back to scans

Will the highest score achieved by a Google Gemini model on Humanity’s Last Exam in 2026 be 50% or higher?

0x1e69d243d667a44cc310c5a3c9ca97cb17a202fb4231c5ddbb74717f5bb4e49e · Science and Technology · 2026-08-21
66%
Agent
70%
Market Price
-4.0%
Edge
56%
Confidence
Volume: 15,118
Spread: 1.6c
Days to resolution: 132
Markets in event: 5
Final Rationale
The Polymarket proxy (~70% YES, uptrending) is the only pricing anchor, but it is thin ($15k volume, 53–97% range) and warrants somewhat less weight than both forecasters gave it. Fundamentals are genuinely favorable: Gemini's best tracked score is 46.44% with a reported 48.4% Deep Think run, rivals have already crossed 50% proving feasibility, and a year-end Gemini 4 / 3.5 Pro release plus competitive pressure to publish a headline HLE number make crossing plausible in ~4.5 months. However, the red team is right that the effective gap from the credited leaderboard value is ~3.5pts rather than ~1.6pts, that recent Gemini progress is non-monotonic/stalled, that Gemini 3.5 Pro is delayed and Gemini 4 was training-only as of Aug 2026, and that resolution-source ambiguity (agi.safe.ai vs. Scale/BenchLM, tools vs. no-tools) plus late-December timing lag are real ways YES fails even if a strong model ships. Netting these, I settle slightly below both forecasts and modestly below the market proxy at 66% YES.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 12$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' listed for any Google Gemini model on the official agi.safe.ai leaderboard, and does that leaderboard include tool-augmented/agentic scores?
  2. What HLE scores have Gemini 3 Pro, Gemini 3 Deep Think, and any later Gemini models reported (no-tools vs with search/code execution)?
  3. What is the observed rate of HLE score improvement from Jan 2025 to now across all frontier models (e.g., ~9% → ~50%+), and what does extrapolation imply for end-2026?
  4. Has any model from any lab already exceeded 50% on HLE, and under what conditions (tools, agentic scaffolds)?
  5. What Google Gemini model releases are expected in 2026 and what benchmark claims have been previewed or leaked?
  6. How frequently is the agi.safe.ai leaderboard updated, and does it lag behind lab-reported scores (resolution risk)?
  7. What is the current Polymarket price for this market and for related thresholds (e.g., 60%, 70%, or other labs' HLE markets)?
Planner reasoning
This is a Polymarket question about whether a Google Gemini model reaches ≥50% on Humanity's Last Exam per the agi.safe.ai leaderboard by end of 2026. The key drivers are the current best Gemini HLE score (Gemini 3 Pro ~37.5% no-tools / ~45% with tools; Deep Think higher), the rate of HLE score improvement across the frontier during 2025, whether the official leaderboard records tool-augmented scores, and Google's expected 2026 release cadence (Gemini 3.5/4, Deep Think updates).
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a Google Gemini model on Humanity’s Last Exam in 2026 be 50% or higher?** - Current price (probability): 70.05% - 7-day price change: +0.30% - 30-day price change: +9.05% - Total volume: $15,118 (USD notional) - Price range: 53.0
polymarket_related OK 2.2s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'Gemini': 0 markets | keyword 'AI benchmark 2026': 0 markets
kalshi_related OK 2.1s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'Gemini': no matches | keyword 'AI model benchmark': ok
claude_news OK 28.5s 12 ## Key Findings on Humanity's Last Exam (HLE) — Gemini Scores - **Current top Gemini score sits below 50%.** The official Scale AI leaderboard (cached ~36 minutes prior to this search) lists gemini-3.1-pro-preview (thinking high) at 46.44±1.96, Calibration Error: 51 , currently the #1 overall mode
claude_news OK 37.3s 15 ## Findings on Google Gemini's HLE Progression & 2026 Outlook - **Gemini 3 Pro (Nov 2025 launch)**: Gemini 3 Pro demonstrates PhD-level reasoning with top scores on Humanity's Last Exam (37.5% without the usage of any tools). This was an 11 percentage-point jump on Humanity's Last Exam, up to 37
gdelt_news OK 175.3s 20 GDELT: 20 articles across 3 queries (lookback=120d). "Gemini Humanity's Last Exam score": error GDELT rate-limited after retries (429) | "Humanity's Last Exam leaderboard record": 10 hits | 'Gemini 3 Deep Think benchmark HLE': 10 hits
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 86.4s 0 ## Key Quantitative Findings **Historical trend data (best Gemini/frontier HLE accuracy, no-tools unless noted):** - Jan 2025 (HLE launch): ~9.1% → Jun 2025 (Gemini 2.5 Pro/Deep Think): ~21.6% → Nov 2025 (Gemini 3 Pro): ~41% → Dec 2025: ~45% - Implied average growth rate: **~3.3 points/month** over
3. Evidence Brief Sonnet · 7050 chars
# Current state The official/tracked HLE leaderboards (Scale AI labs.scale.com, mirrored on Wikipedia) currently show Gemini 3.1 Pro Preview as Google's best-scoring model at **46.44% accuracy (no-tools)**, still below the 50% threshold. Rival labs (Anthropic's Claude Fable 5/Opus 5, OpenAI's GPT-5.6 Sol) have already crossed 50% per some trackers, proving the bar is reachable, but Gemini has not yet done so with ~4.5 months remaining until the Dec 31, 2026 resolution deadline. # Timeline of key events - 2025-01: HLE launches; top model scores ~9% (confirmed, per Wikipedia/code_execution historical data). - 2025-06: Gemini 2.5 Pro/Deep Think reaches ~21.6% (reported, code_execution synthesis). - 2025-11: Gemini 3 Pro launches at 37.5% (no tools), beating GPT-5.1's 26.5% (confirmed, blog.google). - 2026-02: Gemini 3 Deep Think launches at 41.0% (no tools); later reports cite an updated 48.4% (no tools) (reported, Google blog/Remio.ai). - 2026-02-19: Gemini 3.1 Pro reported at ~44.4–46.44%; described as not beating Claude Opus 4.6 (reported, Substack/Texas A&M). - 2026-06 to 2026-08: Gemini 3.5 Pro launch delayed (base-model transition, reasoning shortcomings); interim Flash-tier releases (3.5 Flash, 3.6 Flash, 3.7 Flash) ship without notable HLE claims (reported, TechTimes/Geeky Gadgets/Fonearena). - 2026-08 (recent): Scale AI leaderboard snapshot shows Gemini 3.1 Pro Preview at 46.44% (top Gemini, #1 overall on that board) (confirmed-ish, Scale/Wikipedia). PricePerToken/BenchLM report Claude and OpenAI models above 50–65% (reported, methodology/tool-use unclear). - 2026-08: Gemini 4 confirmed as a training run only; Pichai states ambition for a "larger base model," no release date; rumors point to Nov/Dec 2026 (reported, emergent.sh/AIToolsReview). # Event Will any Google Gemini model reach ≥50% accuracy on the official Humanity's Last Exam leaderboard by Dec 31, 2026? # Outcomes to forecast - Yes - No # Kalshi market anchor **No kalshi_direct price was returned in research** — this is a gap. The only direct pricing data available is from Polymarket (same underlying event, likely mirrored/aggregated): **YES ≈ 70.05%**, up +9.05% over 30 days, range 53.05%–96.85%, volume $15,118. Treat this as the best available consensus proxy in place of a missing Kalshi anchor. # Sub-question answers 1. **Current highest Gemini HLE score / tool inclusion?** Best confirmed Gemini score is 46.44% (Gemini 3.1 Pro Preview, "thinking high," no tools) per Scale AI leaderboard/Wikipedia mirror (claude_news). Resolution source is agi.safe.ai specifically, not directly queried in this research — a gap. 2. **Gemini 3 Pro / Deep Think / later scores?** Gemini 3 Pro: 37.5% (no tools, Nov 2025); Gemini 3 Deep Think: 41.0% initially, later updated to 48.4% (no tools, Feb 2026); Gemini 3.1 Pro: 44.4–46.44% (no tools, varies by tracker). No tool-augmented Gemini scores are cited in sourced material. 3. **HLE growth rate & extrapolation?** ~9% (Jan 2025) → ~46% (Aug 2026), averaging ~2–3 pts/month early on, decelerating recently. Code-execution Monte Carlo modeling gives wide range (55–85%+ under various fits); a discounted judgment estimate lands at ~60–70% probability of Gemini reaching ≥50% by Dec 2026. 4. **Has any model exceeded 50%?** Yes — Claude Fable 5 (55.5%), Claude Opus 5 (54.9%), GPT-5.6 Sol (49.5%) per PricePerToken (Aug 2026); BenchLM shows even higher (Opus 5 at 64.7%), likely reflecting tool-use/methodology differences (unreconciled discrepancy). 5. **2026 Gemini releases/benchmarks previewed?** Gemini 3.5 Pro delayed (base-model rebuild, reasoning issues); interim Flash-tier models (3.5/3.6/3.7 Flash) shipped without major HLE claims; Gemini 4 remains in training only as of Aug 2026, with no confirmed release date (rumored Nov/Dec 2026). 6. **Leaderboard update frequency/lag?** Not explicitly documented; one snapshot was cached "~36 minutes" prior to query, suggesting frequent updates. Multiple third-party trackers (Scale, PricePerToken, BenchLM, Artificial Analysis) show inconsistent numbers for the same models — meaningful resolution-source risk. 7. **Polymarket price for this and related markets?** This exact market: 70.05% YES on Polymarket. A separate "Gemini 3 Predictions" page cites a "50%+" outcome near 73% (claude_news, unverified page). No 60%/70% threshold sibling markets were found via polymarket_related scan (0 matches). # Key facts (high-confidence, factual) 1. [Google blog] Gemini 3 Pro scored 37.5% (no tools) at Nov 2025 launch. 2. [Google blog/Remio.ai] Gemini 3 Deep Think scored 41.0%, later reported at 48.4% (no tools, Feb 2026). 3. [Scale AI/Wikipedia] Gemini 3.1 Pro Preview currently tops Gemini scores at 46.44% (no tools), as of ~Aug 2026. 4. [PricePerToken] Claude Fable 5 (55.5%) and GPT-5.6 Sol (49.5%) have already reached/exceeded 50% on at least one tracker. 5. [emergent.sh] Gemini 4 is in training only as of Aug 2026; no release date confirmed. # Cross-market signals - Kalshi related: no direct Gemini/HLE match found; unrelated markets only. - Polymarket: 70.05% YES on this exact market, uptrending (+9pp/30d); a related "Gemini 3 Predictions" page cites ~73% for a "50%+" outcome (unverified, likely same/related market). - Sportsbook implied: N/A. # Analyst opinions and speculation - Neurosciencenews frames 50% as a real ceiling: "Even the most advanced current models struggle to exceed 50%." - Code-execution quantitative model, after discounting logistic overfit, lands at ~60–70% probability of crossing 50%, with tool-augmented scoring inclusion being the key swing factor. - FutureHouse contamination critique (~30% of chem/bio answers possibly wrong) could distort future scoring dynamics either direction. # Directional lean per outcome - **Yes**: Steady multi-model upward trend (37.5%→41%→48.4%→46.44%); rivals already past 50%, proving feasibility; Gemini 4 possible before year-end; Polymarket pricing at 70%. - **No**: Current best Gemini score (46.44%) still short by ~3.5+ points; Gemini 3.5 Pro delayed with quality issues; Gemini 4 unconfirmed/training-only with no near-term date; recent trajectory shows deceleration/non-monotonicity (46.44% is actually below Deep Think's reported 48.4%). # Gaps / unknowns - No kalshi_direct price was retrieved — primary anchor is missing; relying on Polymarket proxy. - Direct agi.safe.ai leaderboard values not fetched; reliance on Scale AI/Wikipedia mirrors introduces sourcing uncertainty. - Discrepancy between "48.4%" Deep Think claim and "46.44%" current leaderboard top is unreconciled (possibly different scoring runs/methodologies). - Tool-augmented score inclusion on the resolution leaderboard is unconfirmed — critical to the 50% threshold. # Calibration anchors - Polymarket YES price (proxy for missing Kalshi anchor): 70.05%, uptrending. - Precedent: benchmark scores in this AI cycle have shown fast, non-monotonic jumps (9%→46% in ~19 months), but also stalls/plateaus per-model; closing a 3.5pt gap within ~4.5 months is plausible given historical pace.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.55
Yes 74%
No 26%
Gemini's best tracked HLE score is 46.44% (Gemini 3.1 Pro Preview), with Deep Think reported as high as 48.4% — a gap of only ~2–4 points to the 50% threshold with ~4.5 months and likely a Gemini 4 or 3.5 Pro release remaining in 2026. Rival labs (Claude Fable 5 at 55.5%, GPT-5.6 Sol ~49.5%) demonstrate the bar is achievable, creating competitive pressure for Google to publish a headline HLE number above 50%. Counterweights: Gemini 3.5 Pro is delayed with reasoning-quality issues, Gemini 4 was only in training as of Aug 2026, recent progress is non-monotonic, and resolution-source/tool-use methodology ambiguity adds risk. The Polymarket proxy at ~70% YES (uptrending +9pp/30d) is the best consensus anchor absent a Kalshi price; I sit slightly above it given the small remaining gap and typical year-end frontier releases.
gpt-5.5
0.57
Yes 68%
No 32%
Anchoring on the best available market proxy, Polymarket prices YES at about 70% despite no Kalshi-direct price being available. The current best tracked Gemini score of 46.44% points toward YES because the remaining gap is only about 3.5 points, and prior Gemini/HLE progress from 37.5% to the mid/high-40s shows that such gains are plausible within months. Rivals reportedly exceeding 50% point toward YES because the benchmark threshold is demonstrably reachable by frontier labs, but Gemini 3.5 delays, Gemini 4 being only in training, and unreconciled leaderboard/methodology discrepancies point toward NO by raising release and resolution-source risk. I shade slightly below the 70% market proxy because the brief flags non-monotonic recent Gemini scores, no confirmed Gemini 4 release date, and uncertainty over whether the relevant official leaderboard will credit any tool-augmented or updated run before year-end.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters treat the 48.4% Deep Think figure as near-confirmed progress toward 50%, but the brief flags this as an "unreconciled discrepancy" against the 46.44% leaderboard-topping score—if 48.4% is not credited on the resolution source (agi.safe.ai), the effective gap is actually larger (~3.5pts from 46.44%, not ~1.6pts from 48.4%), and neither forecast explicitly discounts for this resolution-source risk in their point estimates. 2. Neither forecaster grapples with the "resolution source" gap seriously: the brief explicitly notes "Direct agi.safe.ai leaderboard values not fetched" and that Scale/Wikipedia/PricePerToken/BenchLM show materially different numbers for the same models—this is a first-order risk for a threshold market that both treat as a minor caveat rather than a core driver of uncertainty. 3. Both anchor heavily on the Polymarket 70% proxy despite explicitly acknowledging it's a substitute for a missing Kalshi price on a low-volume ($15,118) market with a huge 53–97% range, suggesting the "consensus" may be noisy/thin and not deserve the anchoring weight given to it (Forecast 1 goes above it, Forecast 2 slightly below, but both stay within a narrow band near 70%). 4. Neither forecast weighs the "tool-augmented" ambiguity concretely: if tool-use scores end up being the official/resolution metric (as rivals' 50%+ claims may reflect), Gemini's tool-augmented score is completely unknown/unreported, which could swing resolution sharply either direction — this is a binary unknown that should widen uncertainty (wider spread between the two forecasts or lower confidence) rather than be treated as a minor adjustment. 5. The deceleration/non-monotonicity point (46.44% actual top score being below the reported-but-uncredited 48.4% Deep Think) is mentioned by both but not actually reflected quantitatively — a true deceleration signal (37.5%→41%→48.4%→46.44%, i.e., recent regression) arguably warrants more of a discount than the ~5-7pts below Polymarket that both apply, given the specific evidence that growth may have stalled rather than merely slowed. 6. Neither forecast explicitly considers the scenario where Gemini 4 or 3.5 Pro launches very late (e.g., Dec 2026) and reports an interim/preview HLE score too close to year-end to be reliably included on the resolution leaderboard before the Dec 31 cutoff — a timing/logistics risk distinct from "will it score high enough."
6. Aggregator Opus · final calibrated forecast
66% (56% confidence)
Yes: 66% No: 34%
The Polymarket proxy (~70% YES, uptrending) is the only pricing anchor, but it is thin ($15k volume, 53–97% range) and warrants somewhat less weight than both forecasters gave it. Fundamentals are genuinely favorable: Gemini's best tracked score is 46.44% with a reported 48.4% Deep Think run, rivals have already crossed 50% proving feasibility, and a year-end Gemini 4 / 3.5 Pro release plus competitive pressure to publish a headline HLE number make crossing plausible in ~4.5 months. However, the red team is right that the effective gap from the credited leaderboard value is ~3.5pts rather than ~1.6pts, that recent Gemini progress is non-monotonic/stalled, that Gemini 3.5 Pro is delayed and Gemini 4 was training-only as of Aug 2026, and that resolution-source ambiguity (agi.safe.ai vs. Scale/BenchLM, tools vs. no-tools) plus late-December timing lag are real ways YES fails even if a strong model ships. Netting these, I settle slightly below both forecasts and modestly below the market proxy at 66% YES.
Pipeline Timing
Total pipeline time: 290.6s
Per-tool research timings shown in the Research section above.