← Back to scans

Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?

0xd30e3b4e2d6bf4594f44ad982b55435c3add71d3622993b0dee67354defaf598 · Science and Technology · 2026-08-21
47%
Agent
52%
Market Price
-4.5%
Edge
46%
Confidence
Volume: 15,327
Spread: 3.0c
Days to resolution: 132
Markets in event: 5
Final Rationale
OpenAI's documented high-water mark is 49.5% (GPT-5.6 Sol, Aug 2026), leaving a 5.5-point gap with ~4.5 months and one announced-but-unreleased major model (Astra) remaining. The critique's deceleration argument is partly self-refuting: OpenAI moved from 43.1% (Apr) to 49.5% (Jul) on the tracked convention, a ~6.4-point gain in three months, so a further 5.5 points by year-end is well within recent pace if Astra ships on time — but Astra having no announced date is the key downside risk, and a slip into 2027 alone likely resolves No. Convention ambiguity (tool-augmented GPT-5.5 already at 52.2%) plus competitive pressure from Anthropic's 55.5% both tilt bullish and offset the thin-market discount on the 49.5% Polymarket anchor. I land marginally below the anchor at 0.47, reflecting that the trend model's lower blended range (35-45%) and the binary dependence on a single unscheduled release justify a modest lean toward No.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 12$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current highest HLE Accuracy listed on the agi.safe.ai leaderboard for any OpenAI model, and as of what date?
  2. Does the agi.safe.ai leaderboard report text-only/no-tools scores, or does it include tool-augmented/agentic runs (e.g., Deep Research, GPT-5 Pro with search+python), and which convention would resolution follow?
  3. What has been the month-over-month rate of improvement in top OpenAI HLE scores from early 2025 through late 2025/2026 (e.g., o3 ~20%, Deep Research ~26%, GPT-5 ~25-42%)?
  4. What is the current overall SOTA on HLE across all labs (Gemini 3, Grok 4/5, Claude), and how far above 55% or below is it?
  5. How frequently is the agi.safe.ai leaderboard updated, and does it promptly add new OpenAI frontier models?
  6. What OpenAI model releases are expected or already shipped in 2026 (GPT-5.5/GPT-6, o-series successors), and have any published HLE claims near or above 55%?
Planner reasoning
This is a Polymarket question about whether any OpenAI model reaches ≥55% HLE accuracy per the agi.safe.ai leaderboard by end of 2026, so the market price is the primary anchor and the key empirical inputs are the current leaderboard state (which OpenAI model leads, at what score, and whether tool-augmented runs are listed) plus the recent rate of improvement. I'll pull the direct Polymarket price, scan for related markets on both venues, and use news/web search to get the latest HLE numbers and OpenAI model release pipeline, then compute a trend-based extrapolation.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.5s 1 ## This Market's Polymarket Data **Will the highest score achieved by an OpenAI model on Humanity’s Last Exam in 2026 be 55% or higher?** - Current price (probability): 49.50% - 7-day price change: -1.50% - 30-day price change: +20.00% - Total volume: $15,327 (USD notional) - Price range: 29.50% -
polymarket_related OK 3.0s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'OpenAI model benchmark': 0 markets | keyword 'AI benchmark 2026': 0 markets
kalshi_related OK 2.8s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'OpenAI model': ok
claude_news OK 30.5s 6 Here are the key findings on Humanity's Last Exam (HLE) leaderboard status and OpenAI model performance (as of research through mid-August 2026): **Overall leaderboard leaders (top models, not OpenAI):** - On Artificial Analysis's evaluation (which reports both a general leaderboard and an "AA-HLE"
claude_news OK 25.2s 6 ## Key Findings on Humanity's Last Exam (HLE) — 2026 Landscape **OpenAI's current best HLE score (as of Aug 2026):** - As of August 11, 2026, the top-scoring model on Humanity's Last Exam is Claude Fable 5 at 55.5%, followed by Claude Opus 5 at 54.9% and GPT-5.6 Sol at 49.5%. (https://pricepertok
gdelt_news OK 201.4s 20 GDELT: 20 articles across 3 queries (lookback=120d). "Humanity's Last Exam score": error GDELT rate-limited after retries (429) | 'HLE benchmark OpenAI': 10 hits | "GPT-5 Humanity's Last Exam accuracy": 10 hits
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 107.7s 0 ## Quantitative Findings **Data points used (decimal-year, accuracy):** - No-tools track: o1 (2025.03, 9%), o3-mini (2025.08, 13%), o3 (2025.27, 24.9%), GPT-5 (2025.62, 25.3%) - With-tools track: Deep Research (2025.12, 26.6%), GPT-5 w/ tools (2025.62, 42%) **Full-sample log-odds linear fit (all 4
3. Evidence Brief Sonnet · 6178 chars
# Current state As of the most recent snapshot (2026-08-11, pricepertoken.com tracker), OpenAI's best publicly reported HLE score is **GPT-5.6 Sol at 49.5%**, still ~5.5 points below the 55% threshold. Anthropic (Claude Fable 5, 55.5%; Claude Opus 5, 54.9%) has already cleared 55% and currently leads the overall leaderboard — but this question resolves solely on **OpenAI's own highest score**, not overall SOTA. # Timeline of key events - 2025-01/2025-04 (confirmed): o1→o3 released; no-tools HLE scores ~9%→24.9% [Wikipedia/code_execution]. - 2025-08-07 (confirmed): GPT-5 launched; HLE ~25.3% no-tools, ~42% with tools [claude_news, code_execution]. - 2026-02-03 (reported): Gemini 3 Pro Preview leads overall at 37.52% [letsdatascience.com]. - 2026-04-23 (confirmed): OpenAI ships GPT-5.5 (not GPT-6); HLE 43.1% no-tools, 52.2% with tools [venturebeat, rdworldonline]. - 2026-06 (reported): Claude Fable 5 emerges as top "smartest" model per press coverage [timesofindia]. - 2026-07-09/10 (confirmed): GPT-5.6 (Sol/Terra/Luna) reaches general availability [iclarified, heise, gdelt]. - 2026-08-01 (confirmed): OpenAI announces next major model "Astra" — no release date, no HLE score published [felloai.com]. - 2026-08-11 (reported): Leaderboard snapshot shows Claude Fable 5 55.5%, Claude Opus 5 54.9%, **GPT-5.6 Sol 49.5%** — OpenAI's current high-water mark [pricepertoken.com]. # Event Will OpenAI's highest Humanity's Last Exam (HLE) accuracy reach ≥55% by Dec 31, 2026, per the official agi.safe.ai leaderboard (or a published alternative if unavailable)? # Outcomes to forecast Yes / No # Kalshi market anchor No direct kalshi_direct tool output was returned. The polymarket_direct tool (same ticker string) shows current price **49.5%**, down 1.5% over 7 days but up 20 points over 30 days (range 29.5%–65%), volume ~$15.3k. This is the best available cross-market proxy for consensus; treat as the anchor with low confidence given missing native Kalshi data. # Sub-question answers 1. **Current highest OpenAI HLE score & date** — 49.5% (GPT-5.6 Sol), per pricepertoken.com snapshot dated 2026-08-11; an earlier Wikipedia snapshot showed GPT-5.4 Pro at 44.32%. No confirmed timestamp for agi.safe.ai itself was retrieved directly. 2. **Text-only vs tool-augmented convention** — Unclear which convention agi.safe.ai currently uses for its headline score. Third-party trackers show large gaps: e.g., GPT-5.5 scored 43.1% no-tools vs 52.2% with tools [venturebeat/rdworldonline]. If tool-augmented scores count, OpenAI is far closer to 55%. 3. **Rate of improvement** — Highly convention-dependent. No-tools track decelerated sharply (o3 24.9%→GPT-5 25.3% over ~4 months, near-zero logit slope) after an early 2025 sprint; with-tools track rose faster (Deep Research 26.6%→GPT-5 42% in 5 months) but has only 2 data points. GPT-5.5/5.6 no-tools scores (43.1%→~44%) suggest renewed but modest gains through mid-2026. 4. **Overall cross-lab SOTA** — Claude Fable 5 leads at 55.5% (Aug 2026), Claude Opus 5 at 54.9%, Gemini 3 Deep Think at 48.4% (no tools). OpenAI trails the frontier by ~5-6 points on the tracked leaderboard. 5. **Leaderboard update cadence** — Not directly confirmed for agi.safe.ai; third-party mirrors (pricepertoken, artificialanalysis, benchlm) update frequently (within days of model releases), implying reasonably prompt inclusion of new frontier models industry-wide. 6. **Upcoming OpenAI releases** — GPT-5.5 (Apr 2026) and GPT-5.6 (Jul 2026) are incremental "point releases," not GPT-6. OpenAI announced "Astra" (Aug 1, 2026) as its next major model, with solved math/CS problems highlighted but no release date, pricing, or HLE score yet. # Key facts (high-confidence, factual) 1. [pricepertoken.com, 2026-08-11] OpenAI's top score: GPT-5.6 Sol, 49.5%. 2. [venturebeat/rdworldonline] GPT-5.5: 43.1% no-tools, 52.2% with tools. 3. [felloai.com, 2026-08-01] OpenAI's next major model "Astra" unreleased, no HLE data. 4. [pricepertoken.com] Anthropic Claude Fable 5 (55.5%) and Opus 5 (54.9%) currently top overall leaderboard. 5. [Wikipedia] HLE benchmark: 2,500 questions, created by CAIS + Scale AI. # Cross-market signals - Kalshi related: No directly comparable market found; adjacent OpenAI/Anthropic markets (IPO race, US equity stake) show unrelated dynamics. - Polymarket (same ticker): 49.5% YES, +20pts over 30 days — suggests recent bullish momentum, plausibly tied to GPT-5.5/5.6 tool-augmented scores nearing 52%. - No sportsbook analog exists. # Analyst opinions and speculation - Claude synthesis: "OpenAI would need ~5.5+ points gain via Astra or further GPT-5.x release — plausible given pace of releases (7 GPT-5.x point releases in under a year) but not yet demonstrated." - Code-execution quantitative model: no-tools trend implies low-moderate (5–25%) probability of ≥55%; with-tools trend implies moderate-high (40–70%); blended estimate ~35–45% YES, reflecting large convention-dependent uncertainty. # Directional lean per outcome - **Yes**: With-tools GPT-5.5 already at 52.2% (close to threshold); rapid OpenAI release cadence (5.5→5.6→Astra); rising Polymarket price (+20pts/30d) shows market has shifted toward Yes. - **No**: Most recent leaderboard snapshot (Aug 2026) puts OpenAI at only 49.5% under the tracked convention; no-tools trend has plateaued since mid-2025; Astra unreleased with no timeline before year-end; Anthropic/Google currently ahead, no confirmed OpenAI score ≥55% anywhere. # Gaps / unknowns - No confirmed data pulled directly from agi.safe.ai itself (all evidence via third-party mirrors); exact grading convention (tools vs no-tools) used by the official leaderboard is unverified. - No native kalshi_direct market price returned — anchor relies on Polymarket cross-listing. - Astra's release timing/capability and possible HLE score entirely unknown. # Calibration anchors - Polymarket cross-listed price (best available anchor): 49.5% YES. - Quantitative model blended estimate: ~35–45% YES. - Precedent: tool-augmented HLE scores have risen faster (26.6%→52.2% in ~17 months) than no-tools (9%→~44% over same period) — convention choice is the single largest swing factor.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.40
Yes 46%
No 54%
OpenAI's best tracked HLE score as of Aug 2026 is 49.5% (GPT-5.6 Sol), leaving a ~5.5-point gap with roughly 4.5 months remaining. The bull case rests on OpenAI's very fast point-release cadence (5.5 in April, 5.6 in July) plus the announced but unreleased 'Astra' model, and on tool-augmented scoring (GPT-5.5 already at 52.2% with tools) — the convention used by the resolving leaderboard is the single largest uncertainty. The bear case is that the no-tools track has plateaued (24.9%→25.3%→43.1%→~44%) and Anthropic, not OpenAI, currently sits above 55%, with no confirmed OpenAI score anywhere near the threshold. The Polymarket cross-listing (the only anchor available) sits at 49.5% with +20pts over 30 days, while a quantitative trend model blends to ~35-45%; I settle slightly below the market anchor at 0.46.
gpt-5.5
0.52
Yes 44%
No 56%
Anchoring on the best available cross-market proxy, Polymarket has YES at 49.5%, but I shade below it because the latest reported OpenAI high-water mark is only GPT-5.6 Sol at 49.5%, still 5.5 points short with under five months left. The current-score evidence points toward No because no OpenAI model is documented at or above 55%, while the convention uncertainty points toward Yes because tool-augmented GPT-5.5 was already reported at 52.2%. The rate-of-improvement evidence is mixed: no-tools progress looks modest/plateauing, but tool-augmented and point-release cadence make a late Astra or GPT-5.x improvement plausible. Cross-lab SOTA points somewhat toward Yes in principle because Anthropic has already cleared 55%, but upcoming-release evidence points toward No because Astra has no release date or HLE score yet.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters land at ~44-46% Yes despite the brief explicitly noting the quantitative trend model blends to only ~35-45%, and the no-tools track (the likely resolving convention given agi.safe.ai's typical no-tools headline methodology) implies just 5-25% — neither forecaster explains why they weight the higher-uncertainty "convention ambiguity" so heavily rather than defaulting toward the lower no-tools-anchored estimate, which is arguably the more standard reading of "the official agi.safe.ai leaderboard." 2. Neither forecast grapples with the fact that GPT-5.6 Sol (July 2026) only moved the no-tools-adjacent score from ~44% to 49.5% in three months, yet both implicitly assume enough runway remains for another ~5.5-point jump by Astra in under five months — this deceleration-after-acceleration pattern (parallel to the o1→o3→GPT-5 plateau) deserves more skepticism than either analysis gives it. 3. Both forecasters overweight the Polymarket anchor as "the market consensus" despite the brief flagging it as a low-confidence proxy (no native Kalshi price, only $15.3k volume, single cross-listing) — a thinly traded $15k market is a weak basis for anchoring within a few points, and neither forecaster discusses discounting further for illiquidity. 4. Neither addresses the scenario where Astra releases before year-end with a headline HLE score reported using whichever convention is more favorable (labs often cherry-pick tool-augmented scores in launch PR), which could cause a sudden jump past 55% that the smooth trend-extrapolation approach would miss — this "PR-convention arbitrage" tail scenario is glossed over. 5. Both forecasters treat the Anthropic lead (55.5%/54.9%) as merely background context, but it's actually a competitive-pressure signal: OpenAI has historically responded quickly to being overtaken (e.g., o-series sprint in early 2025), which could be a bullish factor for a rushed Astra release/scoring push before Dec 31 that neither fully incorporates. 6. Both converge suspiciously close to each other (0.46 vs 0.44) despite emphasizing different evidence paths, suggesting anchoring on each other's likely output or the market price rather than independently resolving the convention-ambiguity question, which the brief calls "the single largest swing factor" — yet neither commits to a clear view on which convention agi.safe.ai actually uses.
6. Aggregator Opus · final calibrated forecast
47% (46% confidence)
Yes: 47% No: 53%
OpenAI's documented high-water mark is 49.5% (GPT-5.6 Sol, Aug 2026), leaving a 5.5-point gap with ~4.5 months and one announced-but-unreleased major model (Astra) remaining. The critique's deceleration argument is partly self-refuting: OpenAI moved from 43.1% (Apr) to 49.5% (Jul) on the tracked convention, a ~6.4-point gain in three months, so a further 5.5 points by year-end is well within recent pace if Astra ships on time — but Astra having no announced date is the key downside risk, and a slip into 2027 alone likely resolves No. Convention ambiguity (tool-augmented GPT-5.5 already at 52.2%) plus competitive pressure from Anthropic's 55.5% both tilt bullish and offset the thin-market discount on the 49.5% Polymarket anchor. I land marginally below the anchor at 0.47, reflecting that the trend model's lower blended range (35-45%) and the binary dependence on a single unscheduled release justify a modest lean toward No.
Pipeline Timing
Total pipeline time: 296.8s
Per-tool research timings shown in the Research section above.