← Back to scans

Will the highest score achieved by a Google Gemini model on Humanity’s Last Exam in 2026 be 50% or higher?

0x1e69d243d667a44cc310c5a3c9ca97cb17a202fb4231c5ddbb74717f5bb4e49e · Science and Technology · 2026-08-29
70%
Agent
73%
Market Price
-2.9%
Edge
54%
Confidence
Volume: 20,953
Spread: 2.0c
Days to resolution: 124
Markets in event: 5
Final Rationale
Gemini sits at 46.44% on the official leaderboard and 48.4% self-reported (Deep Think, Feb 2026), so only ~1.6-3.6 points separate the current state from the threshold, and Google's annual cadence (Gemini 3 launched Nov 2025) makes a late-2026 frontier release or Deep Think refresh plausible within the remaining ~4 months. The critique is right that mid-2026 deceleration (+2pts in six months) and the six-month failure to reproduce 48.4% on the strict CAIS protocol are genuinely bearish, and that the thin Polymarket price (~$21k, 53-97% range) deserves less than full anchoring weight. Offsetting this, the question asks for the 'highest score achieved by a Google Gemini model,' a framing that plausibly admits Google's own reported figures and higher-effort/tool-inclusive configurations, and rivals already clearing 50-55% on some aggregators shows both feasibility and competitive pressure. Netting the resolution-methodology ambiguity (55-65% no-tools vs 80-85% tools-inclusive) against release-timing and administrative-lag risk, I settle slightly below the market anchor at 70% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 5$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 66% 70% 56%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct claude_news claude_news gdelt_news polymarket_related kalshi_related wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' listed on the official agi.safe.ai leaderboard, and which model holds it?
  2. What HLE accuracy has Google reported for Gemini 3 Pro and Gemini 3 Deep Think (no-tools vs. search/code-execution settings), and which figure does the agi.safe.ai leaderboard use?
  3. How fast has the frontier HLE score risen historically (e.g., ~3% in Jan 2025 → 25-30% mid-2025 → ~37-45% late 2025), i.e., what monthly/annual improvement rate implies crossing 50%?
  4. Is Google expected to release a new frontier model (Gemini 3.5 / 4) in 2026, and on what timeline?
  5. Is the agi.safe.ai leaderboard actively maintained and does it promptly add new Gemini models, or does it lag/omit them (resolution risk)?
  6. What do parallel prediction markets (other thresholds like 40%/60%, or OpenAI/Anthropic equivalents) imply about the probability distribution of 2026 top HLE scores?
Planner reasoning
This is a Polymarket question about whether a Google Gemini model tops 50% HLE accuracy on the official agi.safe.ai leaderboard during 2026, so the market price is the primary anchor and the key empirical inputs are the current leaderboard state, Gemini 3's reported HLE numbers (with/without tools), the historical rate of HLE score improvement, and whether/how frequently the leaderboard is updated. I'll pull the direct market, news on Gemini HLE scores and leaderboard updates, related markets on both venues for other model/threshold variants, and use code execution to extrapolate the improvement trend.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a Google Gemini model on Humanity’s Last Exam in 2026 be 50% or higher?** - Current price (probability): 74.20% - 7-day price change: +3.30% - 30-day price change: -19.80% - Total volume: $20,953 (USD notional) - Price range: 53.
claude_news OK 28.7s 14 ## Key Findings on Humanity's Last Exam (HLE) & Google Gemini Models - **Original Gemini 3 launch (Nov 2025):** Google announced Gemini 3 Pro demonstrates PhD-level reasoning with top scores on Humanity's Last Exam (37.5% without the usage of any tools) and GPQA Diamond (91.9%) . Gemini 3 Deep Thi
claude_news OK 26.0s 11 Here are the key findings on HLE score progression and Google Gemini's trajectory: **Historical progression (Jan 2025 – early 2026):** - At launch (Jan 2025), scores were in the single digits: OpenAI's GPT-4o achieved 3.3% accuracy on Humanity's Last Exam, while Grok-2 achieved 3.8%, Claude 3.5 So
gdelt_news OK 82.5s 20 GDELT: 20 articles across 3 queries (lookback=90d). "Humanity's Last Exam benchmark score": error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28humanity+OR+last+OR+exam+OR+benchmark+OR+score%29+sourcelang%3Aenglish&mode=ArtList&fo
polymarket_related OK 2.1s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Gemini': 0 markets | keyword 'AI benchmark': 0 markets
kalshi_related OK 2.0s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Gemini': no matches
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 88.8s 0 ## Quantitative Findings on Gemini HLE Trajectory → Dec 31 2026 **Data anchors used** (Gemini-family HLE checkpoints, months since Jan 2025): t=0 → ~5% (no-tools)/~9% (tools); t=3 → ~18%/~22% (Gemini 2.5 Pro); t=6 → ~24%/~28% (2.5 Pro update); t=10 (Nov 2025) → 37.5% no-tools / 45.8% tools / ~41% D
3. Evidence Brief Sonnet · 7240 chars
# Current state As of the latest data (~August 2026), the highest verified Google Gemini score on the official CAIS/Scale AI leaderboard is **Gemini 3.1 Pro (thinking high) at 46.44%**, still below the 50% threshold; Google's self-reported Gemini 3 Deep Think figure (48.4%, no-tools, Feb 2026) has not yet appeared as the top Gemini entry on the official leaderboard. Rival labs (Anthropic, OpenAI) have already exceeded 50% on some leaderboards, showing the bar is achievable, but Gemini specifically has not officially cleared it as of now, with ~4-5 months remaining until resolution (Dec 31, 2026). # Timeline of key events - **2025-01**: HLE launches; frontier models score single digits (GPT-4o ~2.7-3.3%, Gemini ~6.2%) (confirmed, techradar.com/intuitionlabs.ai). - **2025-11 (Nov)**: Gemini 3 Pro launches at 37.5% HLE (no tools); Gemini 3 Deep Think at 41.0% (no tools) — Google's own announcement (confirmed, blog.google). - **2026-02-12**: Google announces Gemini 3 Deep Think update reaching 48.4% (no tools) and 84.6% ARC-AGI-2 (confirmed via blog.google, though not yet reflected on third-party leaderboard). - **2026-02-19**: Gemini 3.1 Pro released; model card lists HLE 44.4% (no tools) (confirmed, techcrunch.com/aicerts.ai). - **2026-07/08**: Google ships incremental Gemini 3.1/3.5/3.7 variants (Flash, Transcribe) — no new frontier HLE record reported for these (confirmed, gdelt/blog.google). - **~2026-08 (as of)**: Official Scale/CAIS leaderboard shows Gemini 3.1 Pro (thinking high) at 46.44% — top Gemini score, ranked #1 overall on that leaderboard vs GPT-5.4 Pro at 44.32% (confirmed, labs.scale.com). Independent Artificial Analysis benchmark shows Gemini 3.1 Pro at 44.7% (confirmed, artificialanalysis.ai). Other aggregators (pricepertoken.com, llm-stats.com) show non-Google models (Claude Fable 5, Claude Opus 5, GPT-5.6 Sol) exceeding 50-55%, but these appear to use different/broader evaluation protocols than the official CAIS leaderboard (reported, methodology unclear/possibly tool-assisted). # Event Will any Google Gemini model reach ≥50% HLE accuracy (per agi.safe.ai/CAIS leaderboard) by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No direct kalshi_direct data returned. Best available cross-platform anchor: Polymarket price on this exact ticker = **74.2% YES** (7-day: +3.3%; 30-day: -19.8%; range 53.05%-96.85%; volume $20,953 over 38 days) — a declining-but-recovering trend suggesting the market has cooled somewhat from earlier high confidence but remains net bullish on Yes. # Sub-question answers 1. **Current highest official HLE score/holder?** Gemini 3.1 Pro (thinking high) at 46.44% (±1.96) leads the official Scale/CAIS leaderboard as of ~Aug 2026, ahead of GPT-5.4 Pro at 44.32% (labs.scale.com). 2. **Gemini 3 Pro/Deep Think reported scores, no-tools vs tools?** Gemini 3 Pro: 37.5% no-tools (Nov 2025). Gemini 3 Deep Think: 41.0% no-tools (Nov 2025), later 48.4% no-tools (Feb 2026, self-reported). Gemini 3.1 Pro: 44.4% (model card) vs 46.44% on official leaderboard (thinking-high config) — leaderboard uses a "thinking high"/tool-inclusive-adjacent config that differs slightly from Google's own no-tools model-card number. 3. **Historical growth rate implying 50% crossing?** Rapid climb from ~3-6% (Jan 2025) to 37.5-41% (Nov 2025) to 44-48% (Feb 2026) to 46.44% (Aug 2026) — growth has visibly decelerated in mid-2026 (only ~+2pts from Feb to Aug), consistent with a saturating curve; a free-parameter logistic fit projects a plateau near 44% (no-tools) vs ~68-70% (tools-inclusive) by Dec 2026 (code_execution analysis). 4. **New Gemini model expected in 2026?** Google has shipped multiple incremental updates (3.1, 3.5, 3.7 Flash/Transcribe) through Aug 2026 but no confirmed "Gemini 4" or major new Deep Think record since Feb 2026; no explicit roadmap found in research for a further frontier jump before Dec 2026. 5. **Leaderboard maintenance/lag risk?** Actively maintained — 101 models evaluated as of Aug 2026 update (claude_news); however Google's self-reported 48.4% has not yet appeared as the top leaderboard entry months after announcement, indicating meaningful reporting lag/discrepancy risk between official blog claims and leaderboard-confirmed scores. 6. **Parallel markets/other labs' scores?** Non-Google models (Claude Fable 5 55.5%, Claude Opus 5 54.9%, GPT-5.6 Sol 49.5%) have crossed 50% on some (possibly broader-scope) aggregators, showing the threshold is achievable industry-wide, but on the stricter official CAIS leaderboard all models remain <50% as of Aug 2026, per one source. # Key facts (high-confidence, factual) 1. [labs.scale.com via claude_news] Official leaderboard top Gemini score = 46.44% (Aug 2026), below 50%. 2. [blog.google] Google self-reported Gemini 3 Deep Think 48.4% (no tools, Feb 2026) — not yet confirmed on official leaderboard. 3. [artificialanalysis.ai] Independent benchmark: Gemini 3.1 Pro 44.7%, still <50%. 4. [techradar/intuitionlabs] HLE started near single digits (Jan 2025); Gemini climbed to ~46% by mid-2026 — steep initial trajectory, decelerating. 5. [polymarket_direct] Market-implied probability = 74.2% Yes, down sharply (-19.8%) over 30 days, suggesting growing skepticism. # Cross-market signals - Kalshi related: No direct Gemini/HLE match found; unrelated markets only. - Polymarket: Same-ticker price 74.2% Yes, volatile (53-97% range), moderate volume (~$21k). - Sportsbook implied: N/A. # Analyst opinions and speculation - claude_news bottom line: Google "would need a further model update" to cross 50% before year-end 2026 given current ~46-48% ceiling and deceleration. - code_execution modeling: wide split — tool-augmented scenario ~80-85% likely to hit 50%; no-tools-only interpretation ~55-65%; naive linear extrapolations (implausible, 80-95%) treated as upper-bound noise, not realistic. # Directional lean per outcome - **Yes**: Rapid historical progress (3%→46% in 19 months), competitors already crossing 50% on some leaderboards, multiple Gemini point-releases still expected through year-end, self-reported 48.4% already very close. - **No**: Official leaderboard growth has visibly decelerated (only +2pts in 6 months to Aug 2026), self-reported 48.4% not yet leaderboard-confirmed (resolution source risk), saturating logistic fit projects plateau near 44%, no confirmed roadmap for a game-changing Gemini 4/Deep Think release before Dec 2026. # Gaps / unknowns - No direct Kalshi price feed obtained (relied on identical-ticker Polymarket price). - Unclear whether "tools" vs "no-tools" scores govern official leaderboard resolution — ambiguity could swing probability ~20pts per code_execution analysis. - No confirmed Gemini 4/3.5 major frontier release timeline for H2 2026. - Discrepancy between Google's self-reported 48.4% and leaderboard's 46.44% unresolved — timing/methodology unclear. # Calibration anchors - Polymarket/cross-ticker price: 74.2% Yes (current), down from highs near 97%, low of 53%. - Precedent: HLE scores rose ~3%→38-46% Gemini-specific in ~19 months; last 6 months (Feb-Aug 2026) showed marked deceleration (+2pts), a key bearish signal for continued rapid gains needed to add another ~4-6+ points by Dec 2026.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.50
Yes 74%
No 26%
The official CAIS/Scale leaderboard shows Gemini 3.1 Pro (thinking high) at 46.44%, only ~3.6 points short of the threshold, and Google has already self-reported 48.4% for an updated Gemini 3 Deep Think (Feb 2026), so the gap is small and a single frontier release or Deep Think update would likely clear it. Countervailing: leaderboard progress decelerated sharply (+2pts Feb–Aug 2026), a saturating fit projects a no-tools plateau near ~44-48%, and there is no confirmed Gemini 4/major Deep Think roadmap for H2 2026, plus reporting-lag risk between Google's blog claims and leaderboard confirmation. Rival labs already exceeding 50% on some aggregators shows the bar is technically achievable and creates competitive pressure for Google to push a year-end frontier model. The same-ticker Polymarket price of 74.2% Yes (down from ~97% but recovering) is the best available consensus anchor, and my analysis does not justify a material deviation from it.
gpt-5.6-sol
0.58
Yes 68%
No 32%
The current official score of 46.44% points toward Yes because Google needs only a 3.56-point improvement, while its self-reported 48.4% Deep Think result is already closer still. Historical progress and competitors' 50%+ results show the threshold is technically achievable, but Gemini's roughly two-point gain from February to August indicates substantial deceleration. The absence of a confirmed major release and the official leaderboard's failure to reproduce Google's 48.4% result increase the risk that no qualifying score appears before year-end. I therefore move below the 74.2% Polymarket Yes anchor to 68%, while retaining a Yes majority because one additional frontier update could plausibly clear the threshold.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters anchor heavily on the Polymarket 74.2% price without questioning its reliability: volume is only ~$21k over 38 days and the price swung wildly (53%-97%), suggesting a thin, noisy market that may not deserve near-equal-weighting with the underlying evidence base. 2. Neither forecast adequately grapples with the "tools vs no-tools" ambiguity flagged in the brief as potentially swinging the outcome ~20pts — the code_execution split (55-65% no-tools vs 80-85% tools-inclusive) implies the resolution methodology is a first-order uncertainty that should be explicitly decomposed rather than folded into a single point estimate. 3. Both treat the 48.4% self-reported Deep Think figure as strong evidence of near-crossing, but the brief notes this number has sat unconfirmed on the official leaderboard for ~6 months (Feb to Aug 2026) — this lag itself is evidence the figure may never fully reproduce under the stricter official protocol, which is more bearish than either forecast credits. 4. Neither incorporates a base-rate check on the specific claim "no confirmed Gemini 4/major Deep Think roadmap for H2 2026" — given ~4-5 months remain and Google's cadence has been incremental (Flash/Transcribe variants, not frontier jumps), the No case deserves more independent weight than a mere countervailing bullet. 5. Both converge on 68-74% Yes with minimal daylight between them despite claiming independent reasoning, and neither explores the tail scenario where the official CAIS leaderboard simply never re-scores Gemini's 48.4% claim before Dec 31 (a structural/administrative resolution risk distinct from raw capability), which could push No probability meaningfully higher than either analysis suggests. 6. The deceleration trend (+2pts in 6 months) is mentioned by both but not modeled quantitatively against the ~3.5pt remaining gap — a simple linear extrapolation of the recent (not historical) growth rate would put expected score at ~47-48% by Dec 2026, right at the threshold, implying this should be closer to a coin-flip than a 68-74% Yes.
6. Aggregator Opus · final calibrated forecast
70% (54% confidence)
Yes: 70% No: 30%
Gemini sits at 46.44% on the official leaderboard and 48.4% self-reported (Deep Think, Feb 2026), so only ~1.6-3.6 points separate the current state from the threshold, and Google's annual cadence (Gemini 3 launched Nov 2025) makes a late-2026 frontier release or Deep Think refresh plausible within the remaining ~4 months. The critique is right that mid-2026 deceleration (+2pts in six months) and the six-month failure to reproduce 48.4% on the strict CAIS protocol are genuinely bearish, and that the thin Polymarket price (~$21k, 53-97% range) deserves less than full anchoring weight. Offsetting this, the question asks for the 'highest score achieved by a Google Gemini model,' a framing that plausibly admits Google's own reported figures and higher-effort/tool-inclusive configurations, and rivals already clearing 50-55% on some aggregators shows both feasibility and competitive pressure. Netting the resolution-methodology ambiguity (55-65% no-tools vs 80-85% tools-inclusive) against release-timing and administrative-lag risk, I settle slightly below the market anchor at 70% Yes.
Pipeline Timing
Total pipeline time: 177.8s
Per-tool research timings shown in the Research section above.