← Back to scans

AI model scores ≥ 90% on FrontierMath Benchmark before 2027?

0x4632e96b7010fa4c1474d876f178539f11a284761c08842db5dd2e737cf6341b · Companies · 2026-08-26
84%
Agent
90%
Market Price
-5.5%
Edge
65%
Confidence
Volume: 119,773
Spread: 3.0c
Days to resolution: 126
Markets in event: 1
Final Rationale
The main Tiers 1-3 leaderboard is already reported at ~89.0% (GPT-5.6 Sol) with roughly four months of releases remaining, the post-v2 correction structurally lifted scores, and Epoch's own researcher expects near-term saturation — all pointing strongly to Yes, consistent with the Polymarket anchor at 89.5%. However, the critique correctly identifies underweighted resolution risk: the 87-89% figures come from third-party aggregators rather than Epoch's official leaderboard, Tier 4 v2 still sits at ~83%, the new Open Problems expansion (3/50 solved) could be folded into 'FrontierMath,' and Epoch has a demonstrated habit of ratcheting difficulty as saturation nears. Anchoring uncertainty is also real given the missing Kalshi price and the 83.5% vs 89.5% Polymarket discrepancy. I therefore sit modestly below both prior forecasts and the market anchor, reserving ~16% for a technicality/variant-driven No while keeping capability-driven Yes as the dominant scenario.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 6$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-11 70% 83% 38%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct claude_news claude_news gdelt_news wikipedia polymarket_related kalshi_related code_execution
Sub-questions (Fermi decomposition)
  1. What is the current SOTA score on the FrontierMath benchmark (Tiers 1-3) as reported by Epoch AI, and by which model/date?
  2. What has been the month-over-month/quarterly trajectory of FrontierMath SOTA scores since the benchmark's launch in Nov 2024 (e.g., 2% -> 25% -> 29% -> ~38%)?
  3. Extrapolating the observed improvement rate, what is the probability the SOTA exceeds 90% before Dec 31, 2026?
  4. Which benchmark variant does the market resolve on (full FrontierMath, Tiers 1-3, or Tier 4), and how does that change the required jump?
  5. Are there any announced or imminent frontier model releases (GPT-5.5/6, Gemini 3.5/4, Claude 5, Grok 5) with claimed FrontierMath results near or above 90%?
  6. What has been the pattern of benchmark saturation timelines for prior math benchmarks (MATH, AIME, GPQA) — how long from launch to >90%?
  7. What is the current Polymarket price and price history for this market, and do related AI-capability markets on Kalshi/Polymarket imply a consistent view?
Planner reasoning
This is a capability-forecasting question hinging on the current SOTA score on Epoch AI's FrontierMath benchmark and the observed rate of improvement. The key empirical inputs are the latest verified scores (Tiers 1-3 vs Tier 4), the historical trajectory since late 2024, and whether the 90% threshold is even plausibly reachable in ~13 months. Market price on Polymarket is the primary anchor, with cross-venue and news checks for any recent step-change (e.g., Gemini 3, GPT-5.x, o-series successors).
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.6s 1 ## This Market's Polymarket Data **AI model scores ≥ 90% on FrontierMath Benchmark before 2027?** - Current price (probability): 89.50% - 7-day price change: -0.50% - 30-day price change: +7.50% - Total volume: $119,773 (USD notional) - Price range: 12.50% - 90.50% - Data points: 121 days
claude_news OK 47.3s 16 Based on my research, here is the historical progression and current state of FrontierMath scores: **Benchmark background:** - FrontierMath launched in November 2024, created by Epoch AI with over 60 mathematicians; evaluating six leading language models found all models had very low performance o
claude_news OK 23.8s 7 ## Findings - **Current leader as of July 2026**: GPT-5.6 Sol (OpenAI) tops the FrontierMath (legacy/main) leaderboard at **89%**, ahead of GPT-5.6 Terra (84.9%) and GPT-5.6 Luna (78.6%), out of 17 evaluated models. (https://benchlm.ai/benchmarks/frontierMath, https://llm-stats.com/benchmarks/front
gdelt_news OK 86.0s 20 GDELT: 20 articles across 3 queries (lookback=90d). 'FrontierMath benchmark score': 10 hits | 'Epoch AI FrontierMath': 10 hits | 'AI model math benchmark record': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Read timed out. (read timeout=30)
wikipedia OK 2.0s 3 Fetched 3 Wikipedia entries (1 missing pages).
polymarket_related OK 1.9s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'FrontierMath': 0 markets | keyword 'AI benchmark': 0 markets | keyword "Humanity's Last Exam": 0 markets | keyword 'AGI 2026': 0 markets
kalshi_related OK 1.8s 1 1 related markets / summaries. keyword 'AI benchmark': ok | keyword 'FrontierMath': no matches | keyword 'artificial intelligence model': no matches
code_execution OK 55.6s 0 **Data & setup:** FrontierMath SOTA trajectory (t = months since Nov 2024): t=0 → 2%, t=1 → 25% (o3), t=7 → 29% (mid‑2025), t=12 → 38% (Gemini 3 Pro, late‑2025). Target date Dec 31 2026 ≈ t=25. - **Naive linear extrapolation (all 4 points):** slope ≈ 2.28 pp/month → predicted accuracy at t=25 ≈ **6
3. Evidence Brief Sonnet · 6538 chars
# Current state As of the most recent data (Aug 2026), third-party aggregators report the SOTA on FrontierMath Tiers 1-3 at ~87-89% (GPT-5.5 Pro, Claude Fable 5, GPT-5.6 Sol) and on the harder, error-corrected Tier 4 v2 at ~83% (GPT-5.6 Sol). Epoch AI's own official leaderboard confirmation of a ≥90% score has not been found in this research; most high figures come from secondary/aggregator sites (llm-stats, benchlm, digitalapplied) rather than Epoch AI's primary page. Resolution requires Epoch AI (or credible-reporting consensus) to confirm ≥90% by Dec 31, 2026. # Timeline of key events - 2024-11: FrontierMath launches; SOTA <2% across 6 models (confirmed, arXiv 2411.04872) - 2024-12-20: OpenAI claims o3 scores 25.2% (reported, self-reported by OpenAI) - 2025-04-18: Epoch's independent eval puts o3 at ~10% (confirmed, Epoch) - 2025-04: o4-mini sets record at 17% (±2%) (confirmed, Epoch AI Twitter) - 2025-mid: SOTA ~29% (reported, secondary sources) - 2025-11: Gemini 3 Pro sets record: 38% (Tiers 1-3), 19% (Tier 4) (confirmed, Epoch AI Twitter) - 2025-12/2026-01: GPT-5.2 Pro reaches Tier 4 31%, Tiers 1-3 ~40.7% (reported, Epoch substack + secondary recap) - 2026-05: Epoch review flags fatal errors in ~1/3 of problems (confirmed, Epoch) - 2026-06-12: Epoch releases v2 correcting errors in 42% of problems; dataset now 338 problems (confirmed, Epoch/GDELT) - 2026-06/07 (post-v2): GPT-5.5 Pro (87.7%) and Claude Fable 5 (87.0%) lead Tiers 1-3 per third-party LM Council leaderboard (reported, not Epoch-official) - 2026-07: GPT-5.6 Sol leads main leaderboard at 89.0% across 17 models (reported, llm-stats/benchlm) - 2026-07-31: Epoch launches "Open Problems" expansion (50 unsolved research problems; 3 solved) (confirmed, Epoch) - 2026-08-24: On corrected Tier 4 v2, GPT-5.6 Sol leads at 83% across 47 models (reported, benchlm) # Event Will a SOTA AI model score ≥90% on the FrontierMath benchmark (per Epoch AI) before Jan 1, 2027? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct pull was returned in raw research (gap). Nearest anchor is Polymarket on the identical question: current YES 89.5% per polymarket_direct (30d trend +7.5%, 7d trend -0.5%, range 12.5-90.5%, $119.8k volume, 121 data points). A separate claude_news citation reports Polymarket YES at 83.5% — likely a slightly stale snapshot; treat 89.5% (direct pull) as more current. # Sub-question answers 1. **Current SOTA (Tiers 1-3)** — ~87-89% as of mid/late 2026 (GPT-5.5 Pro 87.7%, Claude Fable 5 87.0%, GPT-5.6 Sol 89.0%), per third-party aggregators (llm-stats, benchlm, digitalapplied); no Epoch-official confirmation found. 2. **Trajectory** — <2% (Nov 2024) → 25.2%/~10% (o3, Dec24/Apr25) → 17% (o4-mini, Apr25) → 38% Tiers1-3/19% Tier4 (Gemini 3 Pro, Nov25) → ~87-89% Tiers1-3 post-v2 correction (mid-2026) → 83% Tier4 v2 (Aug 2026). 3. **Extrapolation** — Code-execution analysis: naive fit (incl. o3 discontinuity) gives P(≥90%)≈20-52%; robust fit excluding that jump gives P≈0% (required pace ~4pp/month vs realized ~1.2-1.8pp/month). However this model is now stale — it predates the June 2026 v2 correction that already pushed scores to ~87-89%, making the extrapolation obsolete. 4. **Benchmark variant** — Market likely resolves on the main/Tiers1-3 leaderboard (already ~89%) rather than the harder Tier 4 v2 (~83%); description says "FrontierMath Exam" generically, primary source Epoch AI, so ambiguity remains on whether Tier 4 counts. 5. **Imminent releases** — GPT-5.6 family (Sol/Terra/Luna) and Claude Fable 5 already released and scoring 78-89%; no confirmed claim of ≥90% yet, but pace suggests imminent crossing. 6. **Prior benchmark saturation** — Not directly answered in research; general pattern (MATH, GPQA) saturates within 1-2 years of launch, consistent with Epoch researcher Greg Burnham's own view that FrontierMath will "probably saturate within the next two years — could be faster" (June 2026). 7. **Polymarket** — 89.5% (direct), up 7.5% over 30 days; no matching related Kalshi/Polymarket markets found (0 matches on keyword scans). # Key facts 1. [arXiv] Nov 2024 launch SOTA <2%. 2. [Epoch AI Twitter] Apr 2025: o4-mini 17%, o3 10%. 3. [Epoch AI Twitter] Nov 2025: Gemini 3 Pro 38% (T1-3)/19% (T4). 4. [Epoch/GDELT] Jun 12 2026: v2 correction fixes errors in 42% of problems, raising scores broadly. 5. [llm-stats/benchlm] Jul 2026: GPT-5.6 Sol 89.0% (T1-3, main leaderboard). 6. [benchlm] Aug 2026: GPT-5.6 Sol 83% on corrected Tier 4 v2. 7. [Epoch AI] Jul 31 2026: New "Open Problems" expansion; only 3/50 solved. 8. [Polymarket direct] Current Polymarket YES = 89.5%. # Cross-market signals - Kalshi related: no direct FrontierMath match found; unrelated ERISA market only. - Polymarket: 89.5% YES, rising 30-day trend, high volume for a niche market ($120k) — strong conviction. - Sportsbook implied: N/A. # Analyst opinions and speculation - Epoch researcher Greg Burnham (Jun 2026): expects saturation "within the next two years — could be faster." - Digitalapplied blog: "today's top scores sit near 88%," implying saturation imminent. - Caveat flagged by claude_news: much of the 87-89% data comes from third-party aggregators, not Epoch's own official page — real Epoch-certified SOTA may lag reported aggregator figures. # Directional lean per outcome - **Yes**: Strong — current aggregator SOTA already ~89% on main leaderboard, trend rising, Polymarket at 89.5%, only ~4 months remain to close and models are already at the threshold; pattern of step-jumps (o3, Gemini 3, GPT-5.5/5.6) suggests another leap plausible. - **No**: Tier 4 (harder problems) still lags at ~83%; if resolution uses full/hardest variant or requires Epoch-official (not aggregator) confirmation, 90% may not be reached; stale trend-model calculations (excluding recent jump) implied near-0% probability, though outdated. # Gaps / unknowns - No kalshi_direct price was retrieved — anchor relies on Polymarket only. - Ambiguity on whether "FrontierMath Exam" resolution uses Tiers1-3 only, full set, or includes Tier 4/Open Problems expansion. - No official Epoch AI leaderboard snapshot confirming exact current SOTA (only third-party sites). - Polymarket price discrepancy (89.5% vs. 83.5% in different tool outputs) unresolved. # Calibration anchors - Polymarket YES 89.5% (primary cross-market anchor, since Kalshi direct missing). - Precedent: MATH/GPQA-style benchmarks saturated within ~1-2 years of launch; FrontierMath already showing similar trajectory 21 months post-launch.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.62
Yes 88%
No 12%
Polymarket on the identical question sits at 89.5% YES with a rising 30-day trend and ~$120k volume, and aggregator leaderboards already put GPT-5.6 Sol at 89.0% on the main Tiers 1-3 FrontierMath leaderboard as of July 2026 — essentially at the threshold with ~4 months of releases remaining. The post-v2 correction (June 2026, 42% of problems fixed) mechanically lifted scores, and Epoch's own Greg Burnham expects saturation possibly within a year. Residual downside risk: much of the 87-89% data is third-party rather than Epoch-official, and the harder Tier 4 v2 still trails at ~83%, so a stricter resolution variant could keep this at No. I anchor near the cross-market price with a small haircut for resolution-source ambiguity.
gpt-5.6-sol
0.68
Yes 85%
No 15%
The latest third-party results point toward Yes because the main Tiers 1-3 leaderboard is reportedly already at 89.0%, leaving only a one-point improvement before the deadline. The trajectory and prior benchmark-saturation pattern also favor Yes, while the stale extrapolation is undermined by June 2026's benchmark correction and subsequent score jump. I move below the 89.5% Polymarket YES anchor because Epoch has not officially confirmed the reported 87-89% scores and the relevant benchmark variant is ambiguous, with Tier 4 v2 still around 83%. Even accounting for those resolution risks, continued model releases over the remaining months make crossing 90% more likely than not by a wide margin.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts converge tightly around the Polymarket anchor (0.88, 0.85) without independently stress-testing the resolution-ambiguity risk highlighted in the brief — specifically, "no Epoch AI (or credible-reporting consensus)" has confirmed ≥90%, only aggregator sites, and neither forecast explores the scenario where Epoch itself never issues an official ≥90% confirmation even if aggregators claim it, which could cause a "No" resolution on a technicality regardless of true model capability. 2. Neither forecast adequately weights the Tier 4 v2 ambiguity as a distinct scenario with real probability mass: the brief notes Tier 4 v2 sits at just 83% and is "harder problems," and if the market resolver (whoever operationally decides, given no Kalshi rules language was even retrieved) leans toward the full/hardest variant, 90% is far less assured — this deserves more than a rhetorical "haircut," perhaps a 15-25% scenario-weighted downside. 3. Both forecasters treat the Polymarket price (89.5%) as a clean anchor, but the brief flags a discrepancy (83.5% in a separate stale snapshot) and zero Kalshi corroboration — meaningful anchoring uncertainty that neither forecast discusses when justifying their small "haircuts" below 89.5%. 4. The "Open Problems" expansion (only 3/50 solved as of July 2026) is mentioned in the brief as a possible resolution variant but is completely ignored by both forecasters — if this expansion feeds into what counts as "FrontierMath" going forward, current trajectory implies scores far below 90%, a scenario neither considered at all. 5. Both forecasts lean on the "step-jump" pattern (o3, Gemini 3, GPT-5.5/5.6) to justify continued acceleration, but this ignores the robust trend-extrapolation finding that excluding the anomalous v2-correction jump, pace is only ~1.2-1.8pp/month — needing ~1pp more by Dec 2026 is plausible, but neither model explicitly reconciles why they dismiss the "robust fit" implying near-0% as fully obsolete rather than partially informative caution. 6. Neither forecast discusses the possibility of benchmark-difficulty ratcheting (Epoch could add harder problems as saturation nears, as they've done historically with v2 and Open Problems) — a structural tendency that could mechanically prevent ≥90% regardless of raw model capability, understating tail "No" risk beyond simple resolution-source ambiguity.
6. Aggregator Opus · final calibrated forecast
84% (65% confidence)
Yes: 84% No: 16%
The main Tiers 1-3 leaderboard is already reported at ~89.0% (GPT-5.6 Sol) with roughly four months of releases remaining, the post-v2 correction structurally lifted scores, and Epoch's own researcher expects near-term saturation — all pointing strongly to Yes, consistent with the Polymarket anchor at 89.5%. However, the critique correctly identifies underweighted resolution risk: the 87-89% figures come from third-party aggregators rather than Epoch's official leaderboard, Tier 4 v2 still sits at ~83%, the new Open Problems expansion (3/50 solved) could be folded into 'FrontierMath,' and Epoch has a demonstrated habit of ratcheting difficulty as saturation nears. Anchoring uncertainty is also real given the missing Kalshi price and the 83.5% vs 89.5% Polymarket discrepancy. I therefore sit modestly below both prior forecasts and the market anchor, reserving ~16% for a technicality/variant-driven No while keeping capability-driven Yes as the dominant scenario.
Pipeline Timing
Total pipeline time: 180.6s
Per-tool research timings shown in the Research section above.