← Back to scans

Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?

0xd30e3b4e2d6bf4594f44ad982b55435c3add71d3622993b0dee67354defaf598 · Science and Technology · 2026-08-30
47%
Agent
50%
Market Price
-3.5%
Edge
50%
Confidence
Volume: 20,553
Spread: 3.0c
Days to resolution: 123
Markets in event: 5
Final Rationale
The governing agi.safe.ai/Scale board shows OpenAI's confirmed best at 44.32% (GPT-5.4 Pro, Mar 2026), ~10.7 points short, and OpenAI is not even the overall leader (Gemini 3.1 Pro, 46.44%). The bullish case rests on unverified third-party figures (GPT-5.6 'Sol' ~49.5%) and on Astra, which as of Aug 2026 has no branding decision, no release, and no published HLE score — a genuine catalyst but far from assured to be benchmarked on the official board before Dec 31. The 55%+ scores circulating belong to Anthropic models, which is evidence of ecosystem progress but not of OpenAI-specific attainment, and OpenAI's own increment pace (44.3→~49.5 over four months, ~1.3 pts/mo) implies roughly landing near, not comfortably above, 55% by year-end. Adding official-verification/lag friction, I sit slightly below the noisy, thin-volume Polymarket proxy of 50.5% rather than above it, unlike both forecasters.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 4$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 47% 52% 46%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest 'HLE Accuracy' value on the official agi.safe.ai leaderboard for any OpenAI model (e.g., GPT-5.x, o-series, Deep Research), and as of what date?
  2. What is the overall state-of-the-art HLE accuracy on that leaderboard across all labs (e.g., Gemini 3 Pro, Claude Opus 4.5), and does the leaderboard report no-tools scores or tool/search-augmented scores?
  3. How fast did the top HLE score rise during 2025 (from ~3-9% in Jan 2025 to what by Dec 2025), i.e., what is the monthly/annual rate of improvement?
  4. Is the agi.safe.ai leaderboard still being actively updated with new frontier model results, and how quickly does it add new OpenAI releases?
  5. What OpenAI model releases are expected/announced for 2026 (GPT-5.5, GPT-6), and what HLE performance have OpenAI or third parties claimed for their newest models?
  6. What do other prediction markets (Polymarket/Kalshi) imply for HLE thresholds at 40%, 50%, 60%, 70% in 2026, and what does the implied distribution say about P(>=55% for OpenAI specifically)?
Planner reasoning
This is a Polymarket AI-benchmark question, so the Polymarket price is the primary anchor, and related markets on other HLE thresholds (e.g., 40%, 50%, 60%) plus Kalshi AI-benchmark markets give a probability curve to triangulate. The key empirical inputs are the current HLE leaderboard values on agi.safe.ai (which reports no-tools scores), the best OpenAI score to date, the historical rate of progress on HLE during 2025, and whether tool-augmented scores count.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?** - Current price (probability): 50.50% - 7-day price change: +3.00% - 30-day price change: +5.50% - Total volume: $20,553 (USD notional) - Price range: 29.50% - 6
polymarket_related OK 2.9s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'OpenAI model benchmark': 0 markets | keyword 'AI benchmark 2026': 0 markets
kalshi_related OK 2.7s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'OpenAI': ok
claude_news OK 28.6s 8 Key findings from research on Humanity's Last Exam (HLE) leaderboard, focused on the official Scale AI/CAIS source (labs.scale.com, linked from agi.safe.ai): - **Official Scale AI leaderboard (no-tools, current as of research)**: gemini-3.1-pro-preview (thinking high) scores 46.44±1.96 with calibr
claude_news OK 22.0s 10 ## Key Findings - **Early 2025 baseline:** When first released in early 2025, state-of-the-art models scored only a few percent – GPT-4o managed just 2.7% and Claude 3.5 Sonnet 4.1% – while subject-matter experts reach around 90% accuracy on the same questions. OpenAI's o1 was the initial top per
gdelt_news OK 83.7s 30 GDELT: 30 articles across 3 queries (lookback=120d). "Humanity's Last Exam benchmark score": 15 hits | 'HLE accuracy OpenAI model': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28hle+OR+accuracy+OR+openai+OR+model%29+sourcela
wikipedia OK 0.1s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 43.0s 0 ## Key Findings **Trend fit on 2025 data (Jan ~9% → mid-year ~26% → Dec ~38-45%)** - Simple linear regression on the no-tools trajectory: slope ≈ **+2.64 pts/month**, R² essentially exact fit through 3 points → naive linear extrapolation to Dec 2026 gives **~70%** (no-tools) and **~83%** if scaled
3. Evidence Brief Sonnet · 7225 chars
# Current state The resolution source (agi.safe.ai → Scale AI/CAIS HLE leaderboard) currently shows OpenAI's best confirmed score as GPT-5.4 Pro at 44.32% (no-tools, ~March 2026), trailing Google's Gemini 3.1 Pro Preview (46.44%) for the overall #1 spot. Third-party aggregators (pricepertoken, llm-stats) report much higher, unverified figures for newer/rumored models (GPT-5.6 Sol ~49.5%, Claude Fable 5 55.5%, Claude Opus 5 54.9–64.7%), but these are NOT confirmed on the official CAIS/Scale leaderboard that governs resolution. # Timeline of key events - 2025-01: HLE launches; SOTA in single digits — o1 8.0%, GPT-4o 2.7%, Claude 3.5 Sonnet 4.1% (confirmed, intuitionlabs/Wikipedia). - 2025-02: OpenAI Deep Research hits 26%, more than doubling prior best (confirmed, Medium/HLE review). - 2025-08-07: GPT-5 launches (confirmed, Wikipedia); mid-2025 HLE score ~25.3% (no-tools), later reports show GPT-5 Pro ~31.6% no-tools / 42.0% with tools (reported). - 2025-08: Grok 4 (xAI) reaches 41–44.4% with tools, briefly framed as tool-assisted leader (reported). - 2025-11: Gemini 3 Pro launches at 37.5% no-tools, becomes new SOTA, surpassing GPT-5 Pro (confirmed, Google blog/press). - 2026-03-05: GPT-5.4 Pro scores 44.32% no-tools, OpenAI's current official-leaderboard best, #2 overall behind Gemini 3.1 Pro at 46.44% (reported, Wikipedia/Scale leaderboard). - 2026-04: GPT-5.5 released; strong on FrontierMath/Terminal-Bench but no confirmed HLE score topping rivals (reported). - 2026-07: GPT-5.6 "Sol" reported at ~49.5% HLE per third-party tracker pricepertoken (reported, low-confidence source). - 2026-08: Third-party trackers show Claude Fable 5 (55.5%) and Claude Opus 5 (54.9–64.7%) as new overall leaders; commentators flag conflict-of-interest concerns re: Meta Muse Spark benchmark claims (rumored/reported, non-official sources). - 2026-08-01: OpenAI unveils "Astra" as its "next major model," undecided whether it will be branded GPT-6 or GPT-5.7; no HLE score published for it (confirmed announcement, no benchmark data). # Event Will any OpenAI model reach ≥55% HLE accuracy on the official agi.safe.ai (Scale AI/CAIS) leaderboard by Dec 31, 2026? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct price was returned in research (tool output missing/empty). Best available cross-market anchor: **Polymarket price 50.5% YES** (up +3% over 7d, +5.5% over 30d; range 29.5–65% over 39 days; volume ~$20.5k) — treat as the closest available consensus proxy pending direct Kalshi confirmation. # Sub-question answers 1. **Current highest OpenAI HLE score on official leaderboard**: GPT-5.4 Pro at 44.32% (no-tools), as of ~March 2026 per Wikipedia's tracking of the Scale AI leaderboard [claude_news]. 2. **SOTA across labs / no-tools vs tools**: Official leaderboard is no-tools; Gemini 3.1 Pro Preview leads at 46.44%, ahead of GPT-5.4 Pro (44.32%) [claude_news]. Tool-augmented scores run higher (GPT-5 Pro 42% with tools, Grok 4 Heavy 44.4%) but the official board reports no-tools figures. 3. **2025 rate of improvement**: SOTA rose from ~8-9% (Jan 2025, o1) to ~37.5% (Nov 2025, Gemini 3 Pro no-tools) to ~44-46% (Mar 2026) — roughly 2.4-2.8 points/month sustained, per code_execution trend fit [code_execution]. 4. **Leaderboard update cadence**: Actively updated; new frontier releases (GPT-5.4, GPT-5.5, GPT-5.6, Gemini 3.1) appear within weeks-to-months of launch, per Wikipedia/Scale tracking [claude_news]. 5. **2026 OpenAI roadmap**: GPT-5.5 (Apr 2026, strong on other benchmarks, no clear HLE leadership); GPT-5.6 "Sol" (Jul 2026, ~49.5% per third-party, unconfirmed officially); "Astra" announced Aug 2026 as next major model, undecided GPT-6 vs GPT-5.7 branding, no HLE score disclosed [claude_news]. 6. **Cross-market implied odds**: Polymarket prices this exact contract at 50.5% YES, trending up. No Kalshi-specific data or other threshold markets (40/50/60/70%) were found in research [polymarket_direct, kalshi_related]. # Key facts (high-confidence, factual) 1. [claude_news/Wikipedia] Official leaderboard best OpenAI score: GPT-5.4 Pro, 44.32% no-tools (~Mar 2026), ~10.7pts below threshold. 2. [claude_news] Overall official-leaderboard leader is Google's Gemini 3.1 Pro Preview at 46.44%, not OpenAI. 3. [claude_news] Third-party (non-official) trackers report OpenAI's newest model (GPT-5.6 Sol) at ~49.5% and rival models (Claude Fable 5, Claude Opus 5) at 54.9–64.7% by Aug 2026 — unverified against the resolution source. 4. [Wikipedia/OpenAI] No GPT-6 confirmed as of Aug 2026; "Astra" announced as next major model with no published HLE score. 5. [code_execution] Historical monthly HLE gains ~2.4-2.8 pts/month across 2025-early 2026 (no-tools). # Cross-market signals - Kalshi related: No direct HLE market found besides this ticker itself; adjacent OpenAI markets (IPO race, US stake) show no HLE-relevant signal. - Polymarket: This exact market prices 50.5% YES, uptrending (+5.5% 30d), moderate volume ($20.5k) — direct proxy since same question. - Sportsbook implied: None available. # Analyst opinions and speculation - code_execution trend-extrapolation model projects Dec-2026 no-tools SOTA of 52-72% depending on deceleration assumption, implying P(≥55%)≈82-95% for "best available" score — but this doesn't isolate OpenAI-specific attainment, and assumes continued linear/logistic growth that official data (GPT-5.4→GPT-5.5→5.6 OpenAI gains) shows may be slowing relative to Google/Anthropic. - claude_news synthesis is more skeptical: OpenAI trails the frontrunner (Anthropic/Google) on official numbers and would need to close a real gap in the remaining months of 2026. # Directional lean per outcome - **Yes**: Rapid 2025 growth trajectory (8%→44% in ~14 months); OpenAI has monthly release cadence (5.4→5.5→5.6→Astra) with each step gaining ground; extrapolation models lean toward eventual ≥55% crossing somewhere in the ecosystem by year-end. - **No**: OpenAI's *own* official-leaderboard best (44.32%) currently trails the overall leader; unverified third-party scores near/above 55% belong to Anthropic (Claude Fable 5/Opus 5), not OpenAI; gap to close for OpenAI specifically is ~10+ points in a market where growth appears to be slowing (5.4%→5.5%→5.6% releases yielding smaller absolute HLE gains); no GPT-6/Astra HLE score exists yet. # Gaps / unknowns - No genuine Kalshi YES price was retrieved for this ticker — brief anchors on Polymarket (50.5%) as substitute; must confirm actual Kalshi price before finalizing. - Third-party tracker figures (pricepertoken, llm-stats) conflict with each other (54.9% vs 64.7% for same model) and aren't validated against the actual agi.safe.ai/Scale leaderboard — high uncertainty on true current OpenAI SOTA post-March 2026. - No confirmed HLE score exists yet for Astra/potential GPT-6, the biggest wildcard for H2 2026. # Calibration anchors - Polymarket YES price (proxy anchor): 50.5%, range 29.5-65% over past ~39 days, trending up. - Precedent: HLE SOTA rose from single digits to ~44-46% in ~14 months (Jan 2025-Mar 2026); reaching 55%+ specifically for OpenAI requires either accelerating past recent deceleration or a major Astra/GPT-6 leap within remaining 2026 window.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.45
Yes 56%
No 44%
The only consensus anchor available is Polymarket at 50.5% YES on this exact question, trending up (+5.5% over 30 days). Official-leaderboard evidence puts OpenAI's confirmed best at GPT-5.4 Pro 44.32% (Mar 2026), ~10.7 points short, and Google leads at 46.44% — but unofficial trackers already put GPT-5.6 'Sol' near 49.5% (Jul 2026) and rivals Claude Fable 5/Opus 5 at 54.9–64.7%, which strongly pressures OpenAI to leap with 'Astra' before year-end. Historical HLE gains of ~2.4–2.8 pts/month over 14 months, plus OpenAI's rapid 5.4→5.5→5.6→Astra cadence and four-plus remaining months, make crossing 55% more likely than not, though resolution depends on the specific agi.safe.ai/Scale board actually posting an OpenAI score ≥55%, and OpenAI's recent increments appear to be trailing Google/Anthropic. I nudge slightly above the Polymarket proxy given the competitive forcing function and the momentum in the price, but keep it near coin-flip given verification/lag risk.
gpt-5.6-sol
0.55
Yes 52%
No 48%
The official 44.32% OpenAI best and 46.44% overall leader point toward No because OpenAI still needs a 10.68-point gain on the governing no-tools leaderboard. However, the historical improvement rate, active leaderboard cadence, and unverified GPT-5.6 estimate near 49.5% point toward Yes by making the remaining gap plausibly achievable. Astra is a major year-end wildcard, while absent official HLE results for GPT-5.5, GPT-5.6, or Astra and signs of slowing gains limit the upward adjustment. With no Kalshi-direct price available, I anchor on Polymarket's 50.5% YES consensus and move slightly upward to 52% due to OpenAI's announced major-model pipeline.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor heavily on Polymarket's 50.5%, but that market itself may be poorly calibrated given thin volume (~$20.5k) and a wide 39-day range (29.5–65%), so treating it as a strong consensus anchor rather than a noisy, low-liquidity proxy risks over-reliance on an unreliable signal. 2. Neither forecast adequately grapples with the "official leaderboard vs. third-party tracker" distinction: the brief explicitly notes GPT-5.6 Sol's ~49.5% is unconfirmed on agi.safe.ai and that resolution requires the *official* board to post ≥55%, yet both nudge probabilities up based on unverified numbers as if they were near-certain to be ratified. 3. Both largely ignore the deceleration signal in the evidence — HLE gains for OpenAI specifically appear to be slowing (5.4→5.5→5.6 yielding diminishing absolute gains) even as the brief's own trend-extrapolation model (52-72%) conflates "best available score across all labs" with "OpenAI's specific score," which the brief explicitly flags as a flaw neither forecaster corrects for. 4. Neither forecast considers the branding/timing risk with Astra: it's "undecided whether it will be branded GPT-6 or GPT-5.7" as of Aug 2026 with "no HLE score published," meaning there's a real chance Astra doesn't launch or get benchmarked before the Dec 31, 2026 close, which should pull probability toward No rather than treating it as a likely catalyst. 5. Both forecasts converge to near-identical probabilities (0.56/0.44 and 0.52/0.48) despite ostensibly independent reasoning, suggesting anchoring bias on the same Polymarket number rather than genuinely stress-testing the OpenAI-specific gap (currently trailing the leader by 2+ points and needing +10.7 points with no confirmed pathway). 6. Neither forecaster addresses the asymmetry that even if Anthropic/Google cross 55% first, this is irrelevant to resolution (question is OpenAI-specific), yet the third-party Claude Fable 5/Opus 5 scores of 54.9–64.7% are used as generic "competitive pressure" evidence for OpenAI's YES case rather than being properly discounted as evidence for a rival, not OpenAI.
6. Aggregator Opus · final calibrated forecast
47% (50% confidence)
Yes: 47% No: 53%
The governing agi.safe.ai/Scale board shows OpenAI's confirmed best at 44.32% (GPT-5.4 Pro, Mar 2026), ~10.7 points short, and OpenAI is not even the overall leader (Gemini 3.1 Pro, 46.44%). The bullish case rests on unverified third-party figures (GPT-5.6 'Sol' ~49.5%) and on Astra, which as of Aug 2026 has no branding decision, no release, and no published HLE score — a genuine catalyst but far from assured to be benchmarked on the official board before Dec 31. The 55%+ scores circulating belong to Anthropic models, which is evidence of ecosystem progress but not of OpenAI-specific attainment, and OpenAI's own increment pace (44.3→~49.5 over four months, ~1.3 pts/mo) implies roughly landing near, not comfortably above, 55% by year-end. Adding official-verification/lag friction, I sit slightly below the noisy, thin-volume Polymarket proxy of 50.5% rather than above it, unlike both forecasters.
Pipeline Timing
Total pipeline time: 186.8s
Per-tool research timings shown in the Research section above.