# Current state
As of the most recent snapshot (2026-08-11, pricepertoken.com tracker), OpenAI's best publicly reported HLE score is **GPT-5.6 Sol at 49.5%**, still ~5.5 points below the 55% threshold. Anthropic (Claude Fable 5, 55.5%; Claude Opus 5, 54.9%) has already cleared 55% and currently leads the overall leaderboard — but this question resolves solely on **OpenAI's own highest score**, not overall SOTA.
# Timeline of key events
- 2025-01/2025-04 (confirmed): o1→o3 released; no-tools HLE scores ~9%→24.9% [Wikipedia/code_execution].
- 2025-08-07 (confirmed): GPT-5 launched; HLE ~25.3% no-tools, ~42% with tools [claude_news, code_execution].
- 2026-02-03 (reported): Gemini 3 Pro Preview leads overall at 37.52% [letsdatascience.com].
- 2026-04-23 (confirmed): OpenAI ships GPT-5.5 (not GPT-6); HLE 43.1% no-tools, 52.2% with tools [venturebeat, rdworldonline].
- 2026-06 (reported): Claude Fable 5 emerges as top "smartest" model per press coverage [timesofindia].
- 2026-07-09/10 (confirmed): GPT-5.6 (Sol/Terra/Luna) reaches general availability [iclarified, heise, gdelt].
- 2026-08-01 (confirmed): OpenAI announces next major model "Astra" — no release date, no HLE score published [felloai.com].
- 2026-08-11 (reported): Leaderboard snapshot shows Claude Fable 5 55.5%, Claude Opus 5 54.9%, **GPT-5.6 Sol 49.5%** — OpenAI's current high-water mark [pricepertoken.com].
# Event
Will OpenAI's highest Humanity's Last Exam (HLE) accuracy reach ≥55% by Dec 31, 2026, per the official agi.safe.ai leaderboard (or a published alternative if unavailable)?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No direct kalshi_direct tool output was returned. The polymarket_direct tool (same ticker string) shows current price **49.5%**, down 1.5% over 7 days but up 20 points over 30 days (range 29.5%–65%), volume ~$15.3k. This is the best available cross-market proxy for consensus; treat as the anchor with low confidence given missing native Kalshi data.
# Sub-question answers
1. **Current highest OpenAI HLE score & date** — 49.5% (GPT-5.6 Sol), per pricepertoken.com snapshot dated 2026-08-11; an earlier Wikipedia snapshot showed GPT-5.4 Pro at 44.32%. No confirmed timestamp for agi.safe.ai itself was retrieved directly.
2. **Text-only vs tool-augmented convention** — Unclear which convention agi.safe.ai currently uses for its headline score. Third-party trackers show large gaps: e.g., GPT-5.5 scored 43.1% no-tools vs 52.2% with tools [venturebeat/rdworldonline]. If tool-augmented scores count, OpenAI is far closer to 55%.
3. **Rate of improvement** — Highly convention-dependent. No-tools track decelerated sharply (o3 24.9%→GPT-5 25.3% over ~4 months, near-zero logit slope) after an early 2025 sprint; with-tools track rose faster (Deep Research 26.6%→GPT-5 42% in 5 months) but has only 2 data points. GPT-5.5/5.6 no-tools scores (43.1%→~44%) suggest renewed but modest gains through mid-2026.
4. **Overall cross-lab SOTA** — Claude Fable 5 leads at 55.5% (Aug 2026), Claude Opus 5 at 54.9%, Gemini 3 Deep Think at 48.4% (no tools). OpenAI trails the frontier by ~5-6 points on the tracked leaderboard.
5. **Leaderboard update cadence** — Not directly confirmed for agi.safe.ai; third-party mirrors (pricepertoken, artificialanalysis, benchlm) update frequently (within days of model releases), implying reasonably prompt inclusion of new frontier models industry-wide.
6. **Upcoming OpenAI releases** — GPT-5.5 (Apr 2026) and GPT-5.6 (Jul 2026) are incremental "point releases," not GPT-6. OpenAI announced "Astra" (Aug 1, 2026) as its next major model, with solved math/CS problems highlighted but no release date, pricing, or HLE score yet.
# Key facts (high-confidence, factual)
1. [pricepertoken.com, 2026-08-11] OpenAI's top score: GPT-5.6 Sol, 49.5%.
2. [venturebeat/rdworldonline] GPT-5.5: 43.1% no-tools, 52.2% with tools.
3. [felloai.com, 2026-08-01] OpenAI's next major model "Astra" unreleased, no HLE data.
4. [pricepertoken.com] Anthropic Claude Fable 5 (55.5%) and Opus 5 (54.9%) currently top overall leaderboard.
5. [Wikipedia] HLE benchmark: 2,500 questions, created by CAIS + Scale AI.
# Cross-market signals
- Kalshi related: No directly comparable market found; adjacent OpenAI/Anthropic markets (IPO race, US equity stake) show unrelated dynamics.
- Polymarket (same ticker): 49.5% YES, +20pts over 30 days — suggests recent bullish momentum, plausibly tied to GPT-5.5/5.6 tool-augmented scores nearing 52%.
- No sportsbook analog exists.
# Analyst opinions and speculation
- Claude synthesis: "OpenAI would need ~5.5+ points gain via Astra or further GPT-5.x release — plausible given pace of releases (7 GPT-5.x point releases in under a year) but not yet demonstrated."
- Code-execution quantitative model: no-tools trend implies low-moderate (5–25%) probability of ≥55%; with-tools trend implies moderate-high (40–70%); blended estimate ~35–45% YES, reflecting large convention-dependent uncertainty.
# Directional lean per outcome
- **Yes**: With-tools GPT-5.5 already at 52.2% (close to threshold); rapid OpenAI release cadence (5.5→5.6→Astra); rising Polymarket price (+20pts/30d) shows market has shifted toward Yes.
- **No**: Most recent leaderboard snapshot (Aug 2026) puts OpenAI at only 49.5% under the tracked convention; no-tools trend has plateaued since mid-2025; Astra unreleased with no timeline before year-end; Anthropic/Google currently ahead, no confirmed OpenAI score ≥55% anywhere.
# Gaps / unknowns
- No confirmed data pulled directly from agi.safe.ai itself (all evidence via third-party mirrors); exact grading convention (tools vs no-tools) used by the official leaderboard is unverified.
- No native kalshi_direct market price returned — anchor relies on Polymarket cross-listing.
- Astra's release timing/capability and possible HLE score entirely unknown.
# Calibration anchors
- Polymarket cross-listed price (best available anchor): 49.5% YES.
- Quantitative model blended estimate: ~35–45% YES.
- Precedent: tool-augmented HLE scores have risen faster (26.6%→52.2% in ~17 months) than no-tools (9%→~44% over same period) — convention choice is the single largest swing factor.