← Back to scans

Will the next Google Gemini Pro model debut with a Humanity’s Last Exam score of 50% or higher?

0x8fca3783aac880990b3c11ffe941e81b1417140429e63dfede4e8b0f6deae839 · Companies · 2026-08-14
53%
Agent
72%
Market Price
-18.5%
Edge
43%
Confidence
Volume: 18,293
Spread: 1.0c
Days to resolution: 139
Markets in event: 4
Final Rationale
Yes requires three things to all go right: (1) the chart isn't effectively frozen by Gemini 3.1 Pro being added at 44.4% (which would mechanically resolve No), (2) a Pro-labeled model actually debuts and is charted before Dec 31, 2026 despite 3.5 Pro's repeated slips, reported hallucination/quality rebuilds, Hassabis's departure, and a roadmap pivot toward 'Gemini 4' of uncertain naming, and (3) that model clears 50% on the no-tools headline protocol, where no frontier model has yet publicly exceeded 50% and Gemini Pro's most recent gen-over-gen gain was only +6.9 pts. Multiplying roughly 0.85-0.9 x 0.75-0.8 x 0.55-0.65 lands near 0.40-0.45, meaningfully below both submitted forecasts and far below the thin Polymarket 71.5% print — which itself just fell 18.5 pts in a week as the delay/leadership news landed, suggesting the tape is trending toward my view. I temper the fully compounded bearish case because Deep Think's 48.4% shows Google's stack is close, a full-generation Pro release would plausibly clear 50%, and labs actively optimize for HLE headlines, so I settle just above even at 0.53 Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 20$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct claude_news claude_news gdelt_news polymarket_related kalshi_related wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What HLE Accuracy does agi.safe.ai currently display for Gemini 3 Pro, and what is the current maximum HLE Accuracy displayed for any model on that site?
  2. How large has the generation-over-generation jump in displayed HLE Accuracy been for Gemini Pro models (2.5 Pro → 3 Pro) and for competitors (GPT-5 → GPT-5.1/5.5, Claude 4.5 → 4.6)?
  3. Is agi.safe.ai's headline 'HLE Accuracy' a no-tools score or does it include search/code execution, and how does that compare to Google's self-reported ~37.5% (no tools) / ~45.8% (with tools) for Gemini 3 Pro?
  4. When is the next Gemini Pro model (e.g., Gemini 3.5 Pro or 4.0 Pro) expected to be released, and how quickly does agi.safe.ai add new models to its results chart?
  5. Has any publicly released frontier model yet exceeded 50% on Humanity's Last Exam under the standardized (non-agentic) evaluation used by agi.safe.ai?
  6. What do other prediction markets (Polymarket/Kalshi) imply about HLE score thresholds and Gemini release timing in 2026?
Planner reasoning
This is a Polymarket question about whether the next Gemini Pro model's debut HLE Accuracy on agi.safe.ai will be ≥50%. Key drivers are (a) what HLE score agi.safe.ai currently displays for Gemini 3 Pro and other frontier models (the site uses a specific no-tools/standardized config that runs lower than vendor-reported tool-augmented numbers), (b) the rate of HLE score improvement per model generation, and (c) whether a new Gemini Pro even appears on the site before Dec 31, 2026. I'll anchor on the market price and then research the leaderboard levels and release timing.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will the next Google Gemini Pro model debut with a Humanity’s Last Exam score of 50% or higher?** - Current price (probability): 71.50% - 7-day price change: -18.50% - 30-day price change: +21.50% - Total volume: $18,293 (USD notional) - Price range: 50.00% - 91.0
claude_news OK 43.1s 13 Here are the key findings on Humanity's Last Exam (HLE) scores relevant to the forecasting question: - **agi.safe.ai site notes headline scores are "no-tools":** Google's own model announcements explicitly frame their agi.safe.ai-referenced HLE scores as no-tools scores — e.g. Gemini 3 Pro: "It de
claude_news OK 29.6s 12 ## Key Findings - **Gemini 3 Pro** (launched Nov 2025) scored 37.5% on Humanity's Last Exam , at the time a state-of-the-art result, while **Gemini 3 Deep Think** scored 41 per cent on Humanity's Last Exam (later benchmarks list up to 48.4% without tools). https://uk.finance.yahoo.com/news/googl
gdelt_news OK 149.6s 20 GDELT: 20 articles across 3 queries (lookback=90d). "Humanity's Last Exam benchmark score record": error GDELT rate-limited after retries (429) | "Gemini 3 Pro Humanity's Last Exam 37.5%": 10 hits | 'Gemini 4 Pro release': 10 hits
polymarket_related OK 3.3s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Gemini': 0 markets | keyword 'Gemini Pro model': 0 markets | keyword 'best AI model': 0 markets
kalshi_related OK 3.2s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'Gemini': no matches | keyword 'AI benchmark': ok
wikipedia OK 0.1s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 44.5s 0 **Key Findings (HLE trend extrapolation for next Gemini Pro model)** - **Data anchors:** Gemini 2.5 Pro ≈ 21% HLE (Mar 25, 2025); Gemini 3 Pro ≈ 37.5% HLE (Nov 18, 2025) — a **+16.5 pt gain over 238 days** (~7.8-month release cadence), giving a raw slope of ~0.069 pts/day (≈2.1 pts/month). - **Tre
3. Evidence Brief Sonnet · 6485 chars
# Current state No Gemini model labeled "Pro" has yet posted an HLE Accuracy of 50%+ on agi.safe.ai's no-tools headline chart. The highest confirmed Gemini Pro figure is Gemini 3.1 Pro at 44.4% (no-tools) / ~51.4% (with tools, not the headline metric). The presumptive "next" qualifying model (Gemini 3.5 Pro) has been delayed repeatedly since its May 2026 I/O announcement and had not shipped as of mid-August 2026; Google has instead been shipping only Flash-tier updates and gesturing toward "Gemini 4." # Timeline of key events - 2025-03-25: Gemini 2.5 Pro ~21% HLE (baseline reference) — reported (code_execution synthesis). - 2025-11-18: Gemini 3 Pro launches, 37.5% HLE no-tools; Gemini 3 Deep Think 41% no-tools — confirmed (Google blog, AOL/Yahoo). - 2026-02-12: Gemini 3 Deep Think (non-Pro variant) reported at 48.4% no-tools, highest Google score to date — confirmed (remio.ai, ekhbary.com). - 2026-02 (approx): Gemini 3.1 Pro reported at 44.4% no-tools / 51.4% with-tools, beating GPT-5.2 (34.5%) but reportedly not beating Claude Opus 4.6 — reported (vertu.com, interestingengineering.com); NOT confirmed whether formally added to agi.safe.ai chart yet. - 2026-05-19: Gemini 3.5 Pro announced at I/O, GA targeted for June; Gemini 3.5 Flash ships instead at 40.2% HLE (regression) — confirmed (techtimes, aitoolsreview). - 2026-06→2026-07-17: Gemini 3.5 Pro GA slips repeatedly (June→July→July 17, then further) — confirmed (findskill.ai, mashable, moneycontrol). - 2026-07-21/22: Google ships three new Flash-tier Gemini models but still no 3.5 Pro; Bloomberg reports internal quality/hallucination issues, base model rebuilt — confirmed (techcrunch, pymnts, moneycontrol). - 2026-07-23: Pichai downplays delay, points to "Gemini 4" and near-monthly release cadence — confirmed (computerworld, moneycontrol). - 2026-08-05: Demis Hassabis steps down as DeepMind CEO amid leadership shakeup — confirmed (fortune.com). - 2026-08-08/13: Gemini 3.5 Pro still in limited Vertex preview, unreleased; Gemini 3.7 Flash ships instead — confirmed (qcode.cc, samaa.tv). # Event Will the next Gemini "Pro"-labeled model added to agi.safe.ai's HLE chart display ≥50% HLE Accuracy? # Outcomes to forecast Yes / No # Kalshi market anchor No Kalshi-direct data was returned for this ticker. The only direct market pricing available is Polymarket for the identical ticker: **71.5% Yes**, down 18.5 pts over 7 days but up 21.5 pts over 30 days; range 50–91%; volume only ~$18.3k across 17 data points — thin, volatile market. Treat as a rough proxy for consensus, not a hard anchor. # Sub-question answers 1. Research could not directly confirm agi.safe.ai's live display for Gemini 3 Pro or the site's current max; press reports (aligned with Google's own no-tools figures) put Gemini 3 Pro at 37.5% and Gemini 3 Deep Think (non-Pro) at up to 48.4%, the apparent current ceiling. 2. Gen-over-gen jumps: Gemini 2.5→3 Pro: +16.5 pts (21%→37.5%); Gemini 3→3.1 Pro: +6.9 pts (37.5%→44.4%) — decelerating. Competitor moves: GPT-5→5.1/5.2 roughly flat/declining (26.5–34.5%); Claude Opus 4.5→4.6 reportedly reached ~53.1% with tools. 3. agi.safe.ai's headline appears to use the no-tools protocol, matching Google's self-reported 37.5% (Gemini 3 Pro) and 44.4% (Gemini 3.1 Pro); with-tools scores (e.g., 45.8%, 51.4%) are higher but not the resolving figure (claude_news). 4. Gemini 3.5 Pro was announced May 2026, delayed multiple times (June→July→ongoing), still unreleased as of Aug 13 2026; Google is emphasizing "Gemini 4" instead, with no firm date (gdelt_news, claude_news). agi.safe.ai's update lag after model release is not documented in research. 5. No frontier model has publicly exceeded 50% no-tools HLE per Google's own reporting; some with-tools scores (Claude Opus 4.6 ~53.1%, Gemini 3.1 Pro ~51.4%) exceed 50% but under a different (non-headline) protocol. Third-party leaderboards (Artificial Analysis 55.5%, BenchLM 64.7%) use divergent, tool-mixed methodologies not comparable to agi.safe.ai's headline metric. 6. Polymarket price 71.5% Yes (volatile, thin). No other Kalshi/Polymarket markets found specifically on HLE thresholds or Gemini release timing. # Key facts (high-confidence, factual) 1. [Google blog] Gemini 3 Pro: 37.5% no-tools HLE (Nov 2025). 2. [vertu.com] Gemini 3.1 Pro: 44.4% no-tools / 51.4% with tools (~Feb 2026). 3. [techcrunch/findskill/qcode.cc] Gemini 3.5 Pro repeatedly delayed, unreleased as of Aug 13, 2026. 4. [fortune.com] Hassabis stepped down as DeepMind CEO Aug 5, 2026 — organizational disruption risk. 5. [computerworld] Pichai now emphasizes "Gemini 4," suggesting 3.5 Pro may be skipped/deprioritized. # Cross-market signals - Kalshi related: no direct Gemini/HLE markets found; unrelated markets only. - Polymarket: 71.5% Yes on this exact ticker, high volatility, low volume/liquidity. - Sportsbook implied: N/A. # Analyst opinions and speculation - code_execution trend extrapolation (2 data points only) projects ~54–67% HLE for a mid-2026 Gemini Pro release, but flags this as likely overstated given benchmark saturation and deceleration already observed (16.5pt→6.9pt jump). - Analysts (aimlapi.com) expect 3.5 Pro to "surpass" prior HLE scores but no leaked/confirmed benchmark exists. # Directional lean per outcome - **Yes**: Historical trend of large jumps (2.5→3 Pro); Deep Think variant already at 48.4%, suggesting Pro tier could follow if given more compute/tools; market still pricing 71.5%. - **No**: Deceleration in gains (3.5x smaller jump 3→3.1 vs 2.5→3); highest confirmed Pro no-tools score is 44.4%, well short of 50%; repeated delays and reported quality/hallucination problems with 3.5 Pro; leadership turmoil (Hassabis exit); Google's own emphasis shifting to "Gemini 4" with no near-term date, risking no qualifying model by Dec 2026 close (resolves No by rule). # Gaps / unknowns - Unclear whether Gemini 3.1 Pro has already been formally added to agi.safe.ai's chart (if so, at 44.4%, this may have already resolved the market No) — critical unresolved fact. - No confirmed release date or benchmark for Gemini 3.5 Pro or Gemini 4. - agi.safe.ai's typical lag between model release and chart addition is undocumented. # Calibration anchors - Polymarket current price: 71.5% Yes (thin liquidity, high volatility — 50-91% range). - Precedent: largest historical single-gen HLE jump for Gemini Pro was +16.5 pts; most recent jump only +6.9 pts, well short of the ~13+ pt jump needed from 44.4%→50%+.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.40
Yes 67%
No 33%
The gap to clear is smaller than the brief's calibration note implies: the current best Pro no-tools score is 44.4% (Gemini 3.1 Pro), so only ~5.6 pts are needed, versus a most-recent gen-over-gen jump of +6.9 pts and an earlier +16.5 pts; the non-Pro Deep Think variant already sits at 48.4% and Gemini 3.1 Pro's with-tools figure is 51.4%, so a next-gen Pro (3.5 Pro after a rebuilt base model, or Gemini 4) plausibly lands above 50%. The main No pathways are (a) no qualifying Pro model being added to agi.safe.ai's chart before the Dec 31, 2026 close — real given repeated 3.5 Pro slips since May 2026, reported hallucination/quality problems, and Hassabis's departure — and (b) a modest, decelerating gain landing in the 45–49% band. I weight ~20-25% on no qualifying release/chart addition and ~80% conditional on release that the score clears 50%, giving roughly 0.65-0.70. This sits just below the thin, highly volatile Polymarket proxy of 71.5% Yes (50–91% range, ~$18k volume), which I discount modestly for liquidity and for the release-timing risk it may underweight.
gpt-5.5
0.46
Yes 55%
No 45%
The rough market anchor is Polymarket at 71.5% Yes, but it is thin and volatile, so I discount it materially. The evidence points toward Yes because Gemini Pro has improved from ~21% to 37.5% to 44.4%, and a non-Pro Deep Think variant has already reached 48.4%, leaving a plausible path to 50% for the next Pro model. The evidence points toward No because the most recent Pro gain was only +6.9 points, the headline agi.safe.ai metric appears to be no-tools rather than with-tools, and Gemini 3.5 Pro has been repeatedly delayed amid quality issues and possible deprioritization toward Gemini 4. A key unresolved risk is whether Gemini 3.1 Pro is or becomes the next chart-added Pro model at only 44.4%, which would resolve No; balancing that against continued model progress, I land modestly below the market but still slightly Yes.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters treat "clearing 50%" and "a qualifying model appearing before close" as largely independent, but the brief's critical unresolved fact—whether Gemini 3.1 Pro (44.4%) has already been formally added to agi.safe.ai's chart—could mechanically resolve the market No regardless of any future model's score; neither forecast assigns explicit probability to this "already resolved/frozen" scenario, which should sharply cap Yes if true. 2. Forecast 1's "only ~5.6 pts needed" framing cherry-picks the smallest possible gap (from 44.4%) while ignoring that the relevant comparison model is unknown/unreleased (Gemini 3.5 Pro or "Gemini 4"), and that the brief explicitly flags decelerating gains (16.5→6.9 pts) as evidence the next jump is more likely to fall in the 45-49% band than clear 50%; an 80% conditional-on-release probability of clearing 50% seems too high given this trend and given reported "quality/hallucination issues" requiring a rebuilt base model. 3. Neither forecast adequately weights the "Gemini 4 instead of 3.5 Pro" branch: if Google skips a Pro-labeled release entirely before Dec 31, 2026 and only ships "Gemini 4" (unclear if it would even carry the "Pro" label), this is a distinct No pathway from "delayed release" that could combine with timeline risk to push total No probability well above what's stated—Pichai's own comments and Hassabis's departure suggest organizational uncertainty about naming/roadmap, not just timing. 4. Both forecasters lean on the Deep Think 48.4% figure as evidence Pro could follow, but Deep Think is explicitly a higher-compute, non-Pro variant (likely using extended reasoning/search), so its score may not transfer to a standard Pro-tier model with different latency/cost constraints—this analogy is weaker than both treat it. 5. Neither critique flags that the Polymarket 71.5% anchor is "down 18.5 pts over 7 days"—a large, fresh negative move—which may reflect market participants pricing in the Aug 2026 delay/Hassabis news more than the forecasts credit; Forecast 1 "discounts modestly" and Forecast 2 "discounts materially" but neither explains why they land above/near a market that just moved sharply against Yes. 6. Both forecasts converge in the 55-67% Yes range without seriously modeling the compounding of three independent risks (chart-timing/frozen-at-44.4% risk, no-Pro-release-by-close risk, and score-shortfall-even-if-released risk), any one of which alone could push No above 50%, suggesting both are overconfident on Yes relative to the brief's cumulative risk factors.
6. Aggregator Opus · final calibrated forecast
53% (43% confidence)
Yes: 53% No: 47%
Yes requires three things to all go right: (1) the chart isn't effectively frozen by Gemini 3.1 Pro being added at 44.4% (which would mechanically resolve No), (2) a Pro-labeled model actually debuts and is charted before Dec 31, 2026 despite 3.5 Pro's repeated slips, reported hallucination/quality rebuilds, Hassabis's departure, and a roadmap pivot toward 'Gemini 4' of uncertain naming, and (3) that model clears 50% on the no-tools headline protocol, where no frontier model has yet publicly exceeded 50% and Gemini Pro's most recent gen-over-gen gain was only +6.9 pts. Multiplying roughly 0.85-0.9 x 0.75-0.8 x 0.55-0.65 lands near 0.40-0.45, meaningfully below both submitted forecasts and far below the thin Polymarket 71.5% print — which itself just fell 18.5 pts in a week as the delay/leadership news landed, suggesting the tape is trending toward my view. I temper the fully compounded bearish case because Deep Think's 48.4% shows Google's stack is close, a full-generation Pro release would plausibly clear 50%, and labs actively optimize for HLE headlines, so I settle just above even at 0.53 Yes.
7. Entry Decision Opus Trader
Side
NO
Position Size
$800
Contracts
2759
Trader Confidence
43%
Trade Rationale
The edge is -18.5pp but this is a direction-agreement setup (bot 53% vs market 71.5%, both above 50%), so the historical base rate demands a higher bar and a smaller size than the raw magnitude suggests. What justifies taking it is that the bearish case rests on concrete, recent, verifiable facts rather than generic pessimism: multiple distinct No pathways (chart freeze if Gemini 3.1 Pro at 44.4% is already charted, no 'Pro'-labeled debut before Dec 31 given 3.5 Pro slips and a possible 'Gemini 4' naming pivot, Hassabis departure, and the empirical fact that no frontier model has publicly cleared 50% no-tools with Gemini Pro's last gen-over-gen gain only +6.9pp). The Devil's Advocate reinforces rather than undermines the NO side on four of five points, and the tape itself has already fallen 18.5pp in a week, indicating the market is drifting toward the forecast rather than away from it. Tempering factors: ensemble spread is 12pp with claude-opus-5 at 67% essentially agreeing with the market, confidence is a middling 0.43, and the 0.53 figure comes from multiplying three plausibly correlated sub-probabilities — the exact miscalibration mode to be wary of.
Allocation Logic
$800 — below the $1000 baseline because this is an agreement-direction edge with a wide ensemble spread and one member near the market price, plus mild correlation with existing AI-benchmark NO positions (Claude HLE score, Anthropic best model); the multiple independent No pathways and the fresh price momentum keep it above the $500 floor.
Entry price: $0.29
Current: $0.34
Status: OPEN
P&L: $151.72
Pipeline Timing
Total pipeline time: 296.7s
Per-tool research timings shown in the Research section above.