# Current state
No Gemini model labeled "Pro" has yet posted an HLE Accuracy of 50%+ on agi.safe.ai's no-tools headline chart. The highest confirmed Gemini Pro figure is Gemini 3.1 Pro at 44.4% (no-tools) / ~51.4% (with tools, not the headline metric). The presumptive "next" qualifying model (Gemini 3.5 Pro) has been delayed repeatedly since its May 2026 I/O announcement and had not shipped as of mid-August 2026; Google has instead been shipping only Flash-tier updates and gesturing toward "Gemini 4."
# Timeline of key events
- 2025-03-25: Gemini 2.5 Pro ~21% HLE (baseline reference) — reported (code_execution synthesis).
- 2025-11-18: Gemini 3 Pro launches, 37.5% HLE no-tools; Gemini 3 Deep Think 41% no-tools — confirmed (Google blog, AOL/Yahoo).
- 2026-02-12: Gemini 3 Deep Think (non-Pro variant) reported at 48.4% no-tools, highest Google score to date — confirmed (remio.ai, ekhbary.com).
- 2026-02 (approx): Gemini 3.1 Pro reported at 44.4% no-tools / 51.4% with-tools, beating GPT-5.2 (34.5%) but reportedly not beating Claude Opus 4.6 — reported (vertu.com, interestingengineering.com); NOT confirmed whether formally added to agi.safe.ai chart yet.
- 2026-05-19: Gemini 3.5 Pro announced at I/O, GA targeted for June; Gemini 3.5 Flash ships instead at 40.2% HLE (regression) — confirmed (techtimes, aitoolsreview).
- 2026-06→2026-07-17: Gemini 3.5 Pro GA slips repeatedly (June→July→July 17, then further) — confirmed (findskill.ai, mashable, moneycontrol).
- 2026-07-21/22: Google ships three new Flash-tier Gemini models but still no 3.5 Pro; Bloomberg reports internal quality/hallucination issues, base model rebuilt — confirmed (techcrunch, pymnts, moneycontrol).
- 2026-07-23: Pichai downplays delay, points to "Gemini 4" and near-monthly release cadence — confirmed (computerworld, moneycontrol).
- 2026-08-05: Demis Hassabis steps down as DeepMind CEO amid leadership shakeup — confirmed (fortune.com).
- 2026-08-08/13: Gemini 3.5 Pro still in limited Vertex preview, unreleased; Gemini 3.7 Flash ships instead — confirmed (qcode.cc, samaa.tv).
# Event
Will the next Gemini "Pro"-labeled model added to agi.safe.ai's HLE chart display ≥50% HLE Accuracy?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct data was returned for this ticker. The only direct market pricing available is Polymarket for the identical ticker: **71.5% Yes**, down 18.5 pts over 7 days but up 21.5 pts over 30 days; range 50–91%; volume only ~$18.3k across 17 data points — thin, volatile market. Treat as a rough proxy for consensus, not a hard anchor.
# Sub-question answers
1. Research could not directly confirm agi.safe.ai's live display for Gemini 3 Pro or the site's current max; press reports (aligned with Google's own no-tools figures) put Gemini 3 Pro at 37.5% and Gemini 3 Deep Think (non-Pro) at up to 48.4%, the apparent current ceiling.
2. Gen-over-gen jumps: Gemini 2.5→3 Pro: +16.5 pts (21%→37.5%); Gemini 3→3.1 Pro: +6.9 pts (37.5%→44.4%) — decelerating. Competitor moves: GPT-5→5.1/5.2 roughly flat/declining (26.5–34.5%); Claude Opus 4.5→4.6 reportedly reached ~53.1% with tools.
3. agi.safe.ai's headline appears to use the no-tools protocol, matching Google's self-reported 37.5% (Gemini 3 Pro) and 44.4% (Gemini 3.1 Pro); with-tools scores (e.g., 45.8%, 51.4%) are higher but not the resolving figure (claude_news).
4. Gemini 3.5 Pro was announced May 2026, delayed multiple times (June→July→ongoing), still unreleased as of Aug 13 2026; Google is emphasizing "Gemini 4" instead, with no firm date (gdelt_news, claude_news). agi.safe.ai's update lag after model release is not documented in research.
5. No frontier model has publicly exceeded 50% no-tools HLE per Google's own reporting; some with-tools scores (Claude Opus 4.6 ~53.1%, Gemini 3.1 Pro ~51.4%) exceed 50% but under a different (non-headline) protocol. Third-party leaderboards (Artificial Analysis 55.5%, BenchLM 64.7%) use divergent, tool-mixed methodologies not comparable to agi.safe.ai's headline metric.
6. Polymarket price 71.5% Yes (volatile, thin). No other Kalshi/Polymarket markets found specifically on HLE thresholds or Gemini release timing.
# Key facts (high-confidence, factual)
1. [Google blog] Gemini 3 Pro: 37.5% no-tools HLE (Nov 2025).
2. [vertu.com] Gemini 3.1 Pro: 44.4% no-tools / 51.4% with tools (~Feb 2026).
3. [techcrunch/findskill/qcode.cc] Gemini 3.5 Pro repeatedly delayed, unreleased as of Aug 13, 2026.
4. [fortune.com] Hassabis stepped down as DeepMind CEO Aug 5, 2026 — organizational disruption risk.
5. [computerworld] Pichai now emphasizes "Gemini 4," suggesting 3.5 Pro may be skipped/deprioritized.
# Cross-market signals
- Kalshi related: no direct Gemini/HLE markets found; unrelated markets only.
- Polymarket: 71.5% Yes on this exact ticker, high volatility, low volume/liquidity.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- code_execution trend extrapolation (2 data points only) projects ~54–67% HLE for a mid-2026 Gemini Pro release, but flags this as likely overstated given benchmark saturation and deceleration already observed (16.5pt→6.9pt jump).
- Analysts (aimlapi.com) expect 3.5 Pro to "surpass" prior HLE scores but no leaked/confirmed benchmark exists.
# Directional lean per outcome
- **Yes**: Historical trend of large jumps (2.5→3 Pro); Deep Think variant already at 48.4%, suggesting Pro tier could follow if given more compute/tools; market still pricing 71.5%.
- **No**: Deceleration in gains (3.5x smaller jump 3→3.1 vs 2.5→3); highest confirmed Pro no-tools score is 44.4%, well short of 50%; repeated delays and reported quality/hallucination problems with 3.5 Pro; leadership turmoil (Hassabis exit); Google's own emphasis shifting to "Gemini 4" with no near-term date, risking no qualifying model by Dec 2026 close (resolves No by rule).
# Gaps / unknowns
- Unclear whether Gemini 3.1 Pro has already been formally added to agi.safe.ai's chart (if so, at 44.4%, this may have already resolved the market No) — critical unresolved fact.
- No confirmed release date or benchmark for Gemini 3.5 Pro or Gemini 4.
- agi.safe.ai's typical lag between model release and chart addition is undocumented.
# Calibration anchors
- Polymarket current price: 71.5% Yes (thin liquidity, high volatility — 50-91% range).
- Precedent: largest historical single-gen HLE jump for Gemini Pro was +16.5 pts; most recent jump only +6.9 pts, well short of the ~13+ pt jump needed from 44.4%→50%+.