# Current state
The resolution source (agi.safe.ai → Scale AI/CAIS HLE leaderboard) currently shows OpenAI's best confirmed score as GPT-5.4 Pro at 44.32% (no-tools, ~March 2026), trailing Google's Gemini 3.1 Pro Preview (46.44%) for the overall #1 spot. Third-party aggregators (pricepertoken, llm-stats) report much higher, unverified figures for newer/rumored models (GPT-5.6 Sol ~49.5%, Claude Fable 5 55.5%, Claude Opus 5 54.9–64.7%), but these are NOT confirmed on the official CAIS/Scale leaderboard that governs resolution.
# Timeline of key events
- 2025-01: HLE launches; SOTA in single digits — o1 8.0%, GPT-4o 2.7%, Claude 3.5 Sonnet 4.1% (confirmed, intuitionlabs/Wikipedia).
- 2025-02: OpenAI Deep Research hits 26%, more than doubling prior best (confirmed, Medium/HLE review).
- 2025-08-07: GPT-5 launches (confirmed, Wikipedia); mid-2025 HLE score ~25.3% (no-tools), later reports show GPT-5 Pro ~31.6% no-tools / 42.0% with tools (reported).
- 2025-08: Grok 4 (xAI) reaches 41–44.4% with tools, briefly framed as tool-assisted leader (reported).
- 2025-11: Gemini 3 Pro launches at 37.5% no-tools, becomes new SOTA, surpassing GPT-5 Pro (confirmed, Google blog/press).
- 2026-03-05: GPT-5.4 Pro scores 44.32% no-tools, OpenAI's current official-leaderboard best, #2 overall behind Gemini 3.1 Pro at 46.44% (reported, Wikipedia/Scale leaderboard).
- 2026-04: GPT-5.5 released; strong on FrontierMath/Terminal-Bench but no confirmed HLE score topping rivals (reported).
- 2026-07: GPT-5.6 "Sol" reported at ~49.5% HLE per third-party tracker pricepertoken (reported, low-confidence source).
- 2026-08: Third-party trackers show Claude Fable 5 (55.5%) and Claude Opus 5 (54.9–64.7%) as new overall leaders; commentators flag conflict-of-interest concerns re: Meta Muse Spark benchmark claims (rumored/reported, non-official sources).
- 2026-08-01: OpenAI unveils "Astra" as its "next major model," undecided whether it will be branded GPT-6 or GPT-5.7; no HLE score published for it (confirmed announcement, no benchmark data).
# Event
Will any OpenAI model reach ≥55% HLE accuracy on the official agi.safe.ai (Scale AI/CAIS) leaderboard by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct price was returned in research (tool output missing/empty). Best available cross-market anchor: **Polymarket price 50.5% YES** (up +3% over 7d, +5.5% over 30d; range 29.5–65% over 39 days; volume ~$20.5k) — treat as the closest available consensus proxy pending direct Kalshi confirmation.
# Sub-question answers
1. **Current highest OpenAI HLE score on official leaderboard**: GPT-5.4 Pro at 44.32% (no-tools), as of ~March 2026 per Wikipedia's tracking of the Scale AI leaderboard [claude_news].
2. **SOTA across labs / no-tools vs tools**: Official leaderboard is no-tools; Gemini 3.1 Pro Preview leads at 46.44%, ahead of GPT-5.4 Pro (44.32%) [claude_news]. Tool-augmented scores run higher (GPT-5 Pro 42% with tools, Grok 4 Heavy 44.4%) but the official board reports no-tools figures.
3. **2025 rate of improvement**: SOTA rose from ~8-9% (Jan 2025, o1) to ~37.5% (Nov 2025, Gemini 3 Pro no-tools) to ~44-46% (Mar 2026) — roughly 2.4-2.8 points/month sustained, per code_execution trend fit [code_execution].
4. **Leaderboard update cadence**: Actively updated; new frontier releases (GPT-5.4, GPT-5.5, GPT-5.6, Gemini 3.1) appear within weeks-to-months of launch, per Wikipedia/Scale tracking [claude_news].
5. **2026 OpenAI roadmap**: GPT-5.5 (Apr 2026, strong on other benchmarks, no clear HLE leadership); GPT-5.6 "Sol" (Jul 2026, ~49.5% per third-party, unconfirmed officially); "Astra" announced Aug 2026 as next major model, undecided GPT-6 vs GPT-5.7 branding, no HLE score disclosed [claude_news].
6. **Cross-market implied odds**: Polymarket prices this exact contract at 50.5% YES, trending up. No Kalshi-specific data or other threshold markets (40/50/60/70%) were found in research [polymarket_direct, kalshi_related].
# Key facts (high-confidence, factual)
1. [claude_news/Wikipedia] Official leaderboard best OpenAI score: GPT-5.4 Pro, 44.32% no-tools (~Mar 2026), ~10.7pts below threshold.
2. [claude_news] Overall official-leaderboard leader is Google's Gemini 3.1 Pro Preview at 46.44%, not OpenAI.
3. [claude_news] Third-party (non-official) trackers report OpenAI's newest model (GPT-5.6 Sol) at ~49.5% and rival models (Claude Fable 5, Claude Opus 5) at 54.9–64.7% by Aug 2026 — unverified against the resolution source.
4. [Wikipedia/OpenAI] No GPT-6 confirmed as of Aug 2026; "Astra" announced as next major model with no published HLE score.
5. [code_execution] Historical monthly HLE gains ~2.4-2.8 pts/month across 2025-early 2026 (no-tools).
# Cross-market signals
- Kalshi related: No direct HLE market found besides this ticker itself; adjacent OpenAI markets (IPO race, US stake) show no HLE-relevant signal.
- Polymarket: This exact market prices 50.5% YES, uptrending (+5.5% 30d), moderate volume ($20.5k) — direct proxy since same question.
- Sportsbook implied: None available.
# Analyst opinions and speculation
- code_execution trend-extrapolation model projects Dec-2026 no-tools SOTA of 52-72% depending on deceleration assumption, implying P(≥55%)≈82-95% for "best available" score — but this doesn't isolate OpenAI-specific attainment, and assumes continued linear/logistic growth that official data (GPT-5.4→GPT-5.5→5.6 OpenAI gains) shows may be slowing relative to Google/Anthropic.
- claude_news synthesis is more skeptical: OpenAI trails the frontrunner (Anthropic/Google) on official numbers and would need to close a real gap in the remaining months of 2026.
# Directional lean per outcome
- **Yes**: Rapid 2025 growth trajectory (8%→44% in ~14 months); OpenAI has monthly release cadence (5.4→5.5→5.6→Astra) with each step gaining ground; extrapolation models lean toward eventual ≥55% crossing somewhere in the ecosystem by year-end.
- **No**: OpenAI's *own* official-leaderboard best (44.32%) currently trails the overall leader; unverified third-party scores near/above 55% belong to Anthropic (Claude Fable 5/Opus 5), not OpenAI; gap to close for OpenAI specifically is ~10+ points in a market where growth appears to be slowing (5.4%→5.5%→5.6% releases yielding smaller absolute HLE gains); no GPT-6/Astra HLE score exists yet.
# Gaps / unknowns
- No genuine Kalshi YES price was retrieved for this ticker — brief anchors on Polymarket (50.5%) as substitute; must confirm actual Kalshi price before finalizing.
- Third-party tracker figures (pricepertoken, llm-stats) conflict with each other (54.9% vs 64.7% for same model) and aren't validated against the actual agi.safe.ai/Scale leaderboard — high uncertainty on true current OpenAI SOTA post-March 2026.
- No confirmed HLE score exists yet for Astra/potential GPT-6, the biggest wildcard for H2 2026.
# Calibration anchors
- Polymarket YES price (proxy anchor): 50.5%, range 29.5-65% over past ~39 days, trending up.
- Precedent: HLE SOTA rose from single digits to ~44-46% in ~14 months (Jan 2025-Mar 2026); reaching 55%+ specifically for OpenAI requires either accelerating past recent deceleration or a major Astra/GPT-6 leap within remaining 2026 window.