← Back to scans

Will the next Google Gemini Pro model debut with a Humanity’s Last Exam score of 50% or higher?

0x8fca3783aac880990b3c11ffe941e81b1417140429e63dfede4e8b0f6deae839 · Companies · 2026-08-24
60%
Agent
70%
Market Price
-10.5%
Edge
52%
Confidence
Volume: 19,588
Spread: 1.0c
Days to resolution: 129
Markets in event: 4
Final Rationale
Two conditions must both hold: a new Pro-labeled Gemini must debut on agi.safe.ai before Dec 31, 2026, and its displayed (historically no-tools) HLE figure must be ≥50%. Timing is the dominant risk — 3.5 Pro has slipped repeatedly and may be skipped, Gemini 4 was only in pre-training as of July 2026, and FutureSearch's median GA (~May 2027) falls after close — so I put roughly 70-75% on any qualifying listing in time. Conditional on a launch, the score bar is modest (+3.6 pts from 46.44%), but the critique is right that a rushed, coding-focused 3.5 Pro or a volatile first-generation Gemini 4 could undershoot, and that Artificial Analysis cross-leaderboard scores (tool-augmented conventions) overstate the case, so ~75-80% conditional is fairer than 85%. That multiplies to ~0.53-0.60. The Polymarket anchor at 70.5% is thin (~$19.6K, 26 points, swinging 50→91) and deserves only partial weight, so I settle modestly below both prior forecasts at 0.60.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 10$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-14 53% 72% 43%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What HLE Accuracy is currently displayed on agi.safe.ai for Gemini 3 Pro, and does the site show no-tools or tool-augmented scores?
  2. What is the highest HLE Accuracy currently displayed on agi.safe.ai for any model (e.g., GPT-5.x, Claude Opus 4.5, Grok 4.x), and has any model crossed 50%?
  3. How much has the top HLE score improved per 6-month interval over 2025-2026, and what jump would Gemini 3 Pro -> next Gemini Pro need to reach 50%?
  4. When is the next Google Gemini Pro model (e.g., Gemini 3.5 Pro / 4.0 Pro) expected to be released, and is there credible reporting on timing before Dec 31, 2026?
  5. How quickly are new frontier models added to agi.safe.ai's HLE results after release (i.e., is a next Gemini Pro likely to appear on the site before the Dec 31, 2026 deadline)?
  6. Are there companion Polymarket/Kalshi markets on other HLE thresholds (e.g., 40%, 45%, 60%) or on other labs' HLE scores that imply a distribution?
Planner reasoning
This is a Polymarket question about whether the next Gemini Pro model's HLE Accuracy on agi.safe.ai will be ≥50%. Key drivers: the current HLE scores displayed on that site (especially Gemini 3 Pro's ~37-38% and frontier leaders), the rate of HLE score improvement per model generation, whether agi.safe.ai reports tool-augmented or no-tools scores, and the expected timing/existence of a next Gemini Pro release before Dec 31, 2026. I'll anchor on the Polymarket price, then gather leaderboard data and release-timeline news, plus cross-venue signals.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.2s 1 ## This Market's Polymarket Data **Will the next Google Gemini Pro model debut with a Humanity’s Last Exam score of 50% or higher?** - Current price (probability): 70.50% - 7-day price change: +0.00% - 30-day price change: +20.50% - Total volume: $19,588 (USD notional) - Price range: 50.00% - 91.00
polymarket_related OK 2.8s 3 Scanned 100 active Polymarket markets, kept 3 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Gemini': 0 markets | keyword 'HLE': 3 markets | keyword 'best AI model': 0 markets
kalshi_related OK 2.4s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'Gemini': no matches | keyword 'AI benchmark': ok
claude_news OK 28.7s 11 Key findings on Humanity's Last Exam (HLE) scores relevant to the "next Gemini Pro" forecast: - **Gemini 3 Pro** (released Nov 2025) debuted with an HLE score of 37.5% without the usage of any tools , per Google's official launch blog. Its "Deep Think" variant scored 41.0% without the use of tool
claude_news OK 24.7s 11 Based on research findings: - **Gemini 3 Pro** (launched Nov 18, 2025) scored **37.5%** on HLE without tools, which Google touted as demonstrating "PhD-level reasoning with top scores on Humanity's Last Exam (37.5% without the usage of any tools)" . (blog.google) - **Gemini 3.1 Pro Preview** (rel
gdelt_news OK 176.2s 20 GDELT: 20 articles across 3 queries (lookback=90d). "Gemini 3 Pro Humanity's Last Exam score": 10 hits | "Humanity's Last Exam benchmark record": error GDELT rate-limited after retries (429) | 'Gemini 4 Pro release': 10 hits
wikipedia OK 0.3s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 48.8s 0 ## Key Findings **Trend fit (no-tools frontier HLE score)** - Linear regression on midpoint scores — Feb 2025: 14.5%, Jul 2025: 30%, Nov 2025: 37.5% (Gemini 3 Pro) — gives slope ≈ **+2.58 pts/month**, intercept ≈12.7%, R²=0.98 (very tight fit, but only 3 points). - Extrapolating to a plausible next
3. Evidence Brief Sonnet · 6405 chars
# Current state As of ~August 2026, the most recent Gemini "Pro" model on the Humanity's Last Exam (HLE) leaderboard is Gemini 3.1 Pro Preview (Feb 2026), scoring 44.4–46.4% — a model the market's own examples (3.2/3.5/4.0 Pro) imply is NOT the "next" qualifying model for this question. The true next Pro-labeled model (Gemini 3.5 Pro or Gemini 4) has not shipped: 3.5 Pro has been repeatedly delayed/possibly skipped, and Gemini 4 is still in pre-training with no confirmed release date, creating real risk that no qualifying model appears before the Dec 31, 2026 deadline at all. # Timeline of key events - 2025-11-18 (confirmed): Gemini 3 Pro launches; HLE 37.5% (no tools) per Google's official blog. - 2026-02-19 (reported): Gemini 3.1 Pro Preview launches; HLE 44.4% per Google, 46.44% per CAIS/agi.safe.ai-sourced Wikipedia table (GPT-5.4 Pro 44.32%, Muse Spark 40.56% same leaderboard). - 2026-06 to 2026-08 (reported): Google ships multiple Flash-tier models (3.5/3.6/3.7 Flash) but 3.5 Pro repeatedly misses target windows (June, mid-July, early August). - 2026-07-17/18/21/23 (reported): Multiple outlets (TechCrunch, Business Insider, Moneycontrol) confirm Gemini 3.5 Pro delay tied to coding-performance issues; Pichai publicly redirects attention to Gemini 4 and promises faster cadence. - 2026-07-21 (reported): Gemini 4 confirmed in pre-training only (kie.ai); no benchmarks, no date. - 2026-08-05 (reported): Google DeepMind reorg (Hassabis to chairman; Jeff Dean heads "Discovery Loop"); geeky-gadgets reports DeepMind may pivot away from 3.5 Pro toward Gemini 4. - 2026-08-13/14 (confirmed): Gemini 3.7 Flash released — non-Pro, does not qualify. - ~2026-08 (market data): Polymarket price for this exact market at 70.5%, up from a 50% floor over the trailing 30 days. # Event Will the next Google Gemini "Pro"-labeled model to appear on agi.safe.ai's HLE leaderboard (after Gemini 3.1 Pro Preview) debut with a displayed HLE Accuracy ≥50%, by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No distinct kalshi_direct price was returned; the only direct price data available for this exact ticker came via polymarket_direct: **YES = 70.5%**, 7-day change flat, 30-day change **+20.5 pts** (rising from a 50% floor to a 91% high over 26 data points), volume ~$19.6K. Treat this as the best available consensus anchor; true Kalshi order-book price is a data gap. # Sub-question answers 1. Gemini 3 Pro is listed at 37.5% (no-tools) on Google's own materials and ~37.52% on the CAIS leaderboard; a tool/search-augmented variant reportedly scored ~45.8%, but the site's headline "HLE Accuracy" card figure appears to be the no-tools number. 2. Highest score currently on agi.safe.ai's leaderboard is Gemini 3.1 Pro Preview at 46.44% (Wikipedia/CAIS mirror); GPT-5.4 Pro 44.32%, Muse Spark 40.56%. No model has crossed 50% on the actual resolution source as of mid-Aug 2026. [Note: third-party Artificial Analysis leaderboard shows some models >50% (Claude Fable 5 ~53-55%, GPT-5.6 Sol ~49.5%), but that is a different site, not the resolution source.] 3. Trend: ~14.5% (Feb'25) → 30% (Jul'25) → 37.5% (Nov'25, Gemini 3 Pro) → 44.4-46.4% (Feb'26, 3.1 Pro) — roughly +7-9 pts per ~3-month Pro iteration (~+2.6 pts/month regression fit). Only ~+3.6 pts needed from 3.1 Pro's 46.44% to reach 50%. 4. Gemini 3.5 Pro has been delayed through Aug 2026 and may be skipped; Gemini 4 is in pre-training only (confirmed July 2026), no date, with FutureSearch's model forecasting median GA ~May 2027 (past close), though some industry chatter clusters on late-2026. 5. Prior Gemini Pro releases (3 Pro, 3.1 Pro) appeared on agi.safe.ai promptly (days/weeks) after launch, so if a new Pro model ships before ~mid-Dec 2026 it should appear in time to resolve. 6. No companion Kalshi/Polymarket HLE-threshold or other-lab markets were found; keyword searches returned unrelated esports/SCOTUS markets only. # Key facts (high-confidence, factual) 1. [Google blog] Gemini 3 Pro: 37.5% HLE, no tools (Nov 2025). 2. [Google/IE/Wikipedia] Gemini 3.1 Pro Preview: 44.4% (Google) / 46.44% (CAIS leaderboard) (Feb 2026). 3. [TechCrunch, Business Insider, Moneycontrol] Gemini 3.5 Pro repeatedly delayed through Aug 2026; only Flash-tier models shipped in interim. 4. [kie.ai] Gemini 4 confirmed in pre-training as of July 2026, no benchmarks/date. 5. [FutureSearch] Median forecast GA for Gemini 4: May 2027 (after market close). # Cross-market signals - Polymarket (this market): YES 70.5%, rising sharply over past 30 days. - Kalshi related: no Gemini/HLE-specific matches; unrelated markets only. - No sportsbook or other AI-benchmark markets found. # Analyst opinions and speculation - Code-execution trend extrapolation: no-tools P(≥50%) ≈17-84% depending on release timing (centered ~38% for a ~Mar 2026 release, but real timing looks much later given delays); tool-augmented P(≥50%) ≈91-99% if a tool-boosted number is ever headlined (unlikely per resolution rules, which track the card's stated "HLE Accuracy," historically the no-tools figure). - geeky-gadgets/computerworld: DeepMind may skip 3.5 Pro and pivot to Gemini 4 as flagship "next Pro." # Directional lean per outcome - **Yes**: Score trajectory needs only ~+3.6 pts from 46.44% baseline, a smaller jump than the last two increments (+7-9 pts each); Gemini 4 rumored to be "significantly larger," suggesting a bigger capability jump. Polymarket already prices 70.5%. - **No**: Substantial release-timing risk — 3.5 Pro delayed/possibly skipped, Gemini 4 still pre-training with no confirmed date and some forecasts (FutureSearch) placing GA after the market's Dec 31, 2026 close, which would resolve "No" outright regardless of score. # Gaps / unknowns - No confirmed kalshi_direct price; relying on Polymarket cross-platform data for the same ticker. - Whether agi.safe.ai's displayed figure will be no-tools or tool-augmented for the next model is not explicitly confirmed. - Exact scoring convention/resolution nuance if Gemini 4 is not labeled "Pro" (could be moot per description's flexible naming) is unresolved. # Calibration anchors - Cross-platform (Polymarket) current YES price: 70.5% (anchor). - Precedent: HLE scores have risen ~7-9 pts per successive Gemini Pro release over ~3 months; only ~7-10 months of runway remain before close, with the presumed next release facing significant delay risk.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.47
Yes 65%
No 35%
Two hurdles must both clear: a new Pro-labeled Gemini must appear on agi.safe.ai before Dec 31, 2026, and its displayed HLE figure must be ≥50%. On score, the bar is modest — 3.1 Pro already sits at 46.4% (Feb 2026) and the trend is ~+7-9 pts per Pro iteration, with third-party leaderboards already showing rival models near or above 50%, so conditional on a launch I put ~80-85% on ≥50%. The larger risk is timing: 3.5 Pro has slipped repeatedly since June and may be skipped, Gemini 4 was only in pre-training as of July 2026, and FutureSearch's median GA (~May 2027) falls after close — I put ~75% on a qualifying Pro shipping and being listed in time, given competitive pressure and Pichai's faster-cadence promise. That yields ~0.62, which I nudge toward the 70.5% Polymarket anchor to ~0.65.
gpt-5.5
0.56
Yes 63%
No 37%
Anchoring on the only direct consensus price, Polymarket YES at 70.5%, I shade lower because the brief highlights substantial timing risk not just score risk. The score evidence points toward Yes: Gemini Pro rose from 37.5% to 44.4–46.4%, and only about +3.6 points are needed for the next Pro-labeled model to clear 50%, a smaller jump than recent iterations. The release-timing evidence points toward No: 3.5 Pro has been repeatedly delayed or possibly skipped, Gemini 4 was still in pre-training in July 2026, and FutureSearch’s median GA forecast is May 2027, after close. If a qualifying Pro model ships by mid-December it should likely appear on agi.safe.ai promptly, so the main uncertainty is whether any qualifying model debuts in time rather than whether its HLE score reaches 50%.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters converge near 0.63-0.65, essentially just shading toward the Polymarket anchor, but neither explicitly stresses that Polymarket is a thin, likely retail-dominated market (~$19.6K volume, 26 data points) that swung from a 50% floor to 91% high in 30 days — this is weak evidence for a "consensus" anchor and shouldn't be weighted as heavily as a liquid market. 2. Forecast 1's implied "conditional on launch" P(≥50%) of 80-85% is inconsistent with its own math: 0.75 (timing) × 0.80-0.85 (score) = 0.60-0.64, yet it rounds to 0.62 then nudges to 0.65 toward the anchor — this "nudge toward anchor" step is somewhat circular reasoning rather than independent evidence. 3. Neither forecaster grapples with the base-rate risk that if Gemini 4 is the qualifying model and is "significantly larger" (per the brief's speculation), a *first* release of a genuinely new architecture could just as easily undershoot the extrapolated trend as overshoot it — new model generations often show volatile, non-monotonic benchmark jumps rather than smooth continuation of the prior scaling line. 4. Both give substantial weight to third-party leaderboards (Artificial Analysis showing >50% for other labs) as evidence the bar is "achievable," but the brief explicitly warns the resolution source is agi.safe.ai's no-tools figure, and Gemini's own historical convention (37.5% no-tools vs 45.8% tool-augmented) means the qualifying number could land lower than analogous third-party site scores — this cross-leaderboard comparison risk is underexplored. 5. Neither forecast addresses the ambiguous case where Gemini 3.5 Pro ships as a genuine "Pro" release before Gemini 4 — the brief notes 3.5 Pro was optimized to fix "coding-performance issues," not necessarily reasoning/HLE performance, so if 3.5 Pro is the qualifying model (rather than Gemini 4), the score jump could be much smaller than the +7-9pt trend, undermining the "only need +3.6pts" framing that both forecasters lean on. 6. Both treat "timing" and "score" as roughly independent multiplicative factors, but they are correlated: a rushed, delayed release under competitive pressure (Pichai's "faster cadence" promise) is arguably more likely to ship with a modest score bump (to hit a deadline) than a fully-baked large jump, meaning P(Yes | late 2026 launch) may be lower than the unconditional trend-based estimate both use.
6. Aggregator Opus · final calibrated forecast
60% (52% confidence)
Yes: 60% No: 40%
Two conditions must both hold: a new Pro-labeled Gemini must debut on agi.safe.ai before Dec 31, 2026, and its displayed (historically no-tools) HLE figure must be ≥50%. Timing is the dominant risk — 3.5 Pro has slipped repeatedly and may be skipped, Gemini 4 was only in pre-training as of July 2026, and FutureSearch's median GA (~May 2027) falls after close — so I put roughly 70-75% on any qualifying listing in time. Conditional on a launch, the score bar is modest (+3.6 pts from 46.44%), but the critique is right that a rushed, coding-focused 3.5 Pro or a volatile first-generation Gemini 4 could undershoot, and that Artificial Analysis cross-leaderboard scores (tool-augmented conventions) overstate the case, so ~75-80% conditional is fairer than 85%. That multiplies to ~0.53-0.60. The Polymarket anchor at 70.5% is thin (~$19.6K, 26 points, swinging 50→91) and deserves only partial weight, so I settle modestly below both prior forecasts at 0.60.
Pipeline Timing
Total pipeline time: 315.3s
Per-tool research timings shown in the Research section above.