# Current state
As of the most recent data (Aug 2026), third-party aggregators report the SOTA on FrontierMath Tiers 1-3 at ~87-89% (GPT-5.5 Pro, Claude Fable 5, GPT-5.6 Sol) and on the harder, error-corrected Tier 4 v2 at ~83% (GPT-5.6 Sol). Epoch AI's own official leaderboard confirmation of a ≥90% score has not been found in this research; most high figures come from secondary/aggregator sites (llm-stats, benchlm, digitalapplied) rather than Epoch AI's primary page. Resolution requires Epoch AI (or credible-reporting consensus) to confirm ≥90% by Dec 31, 2026.
# Timeline of key events
- 2024-11: FrontierMath launches; SOTA <2% across 6 models (confirmed, arXiv 2411.04872)
- 2024-12-20: OpenAI claims o3 scores 25.2% (reported, self-reported by OpenAI)
- 2025-04-18: Epoch's independent eval puts o3 at ~10% (confirmed, Epoch)
- 2025-04: o4-mini sets record at 17% (±2%) (confirmed, Epoch AI Twitter)
- 2025-mid: SOTA ~29% (reported, secondary sources)
- 2025-11: Gemini 3 Pro sets record: 38% (Tiers 1-3), 19% (Tier 4) (confirmed, Epoch AI Twitter)
- 2025-12/2026-01: GPT-5.2 Pro reaches Tier 4 31%, Tiers 1-3 ~40.7% (reported, Epoch substack + secondary recap)
- 2026-05: Epoch review flags fatal errors in ~1/3 of problems (confirmed, Epoch)
- 2026-06-12: Epoch releases v2 correcting errors in 42% of problems; dataset now 338 problems (confirmed, Epoch/GDELT)
- 2026-06/07 (post-v2): GPT-5.5 Pro (87.7%) and Claude Fable 5 (87.0%) lead Tiers 1-3 per third-party LM Council leaderboard (reported, not Epoch-official)
- 2026-07: GPT-5.6 Sol leads main leaderboard at 89.0% across 17 models (reported, llm-stats/benchlm)
- 2026-07-31: Epoch launches "Open Problems" expansion (50 unsolved research problems; 3 solved) (confirmed, Epoch)
- 2026-08-24: On corrected Tier 4 v2, GPT-5.6 Sol leads at 83% across 47 models (reported, benchlm)
# Event
Will a SOTA AI model score ≥90% on the FrontierMath benchmark (per Epoch AI) before Jan 1, 2027?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct pull was returned in raw research (gap). Nearest anchor is Polymarket on the identical question: current YES 89.5% per polymarket_direct (30d trend +7.5%, 7d trend -0.5%, range 12.5-90.5%, $119.8k volume, 121 data points). A separate claude_news citation reports Polymarket YES at 83.5% — likely a slightly stale snapshot; treat 89.5% (direct pull) as more current.
# Sub-question answers
1. **Current SOTA (Tiers 1-3)** — ~87-89% as of mid/late 2026 (GPT-5.5 Pro 87.7%, Claude Fable 5 87.0%, GPT-5.6 Sol 89.0%), per third-party aggregators (llm-stats, benchlm, digitalapplied); no Epoch-official confirmation found.
2. **Trajectory** — <2% (Nov 2024) → 25.2%/~10% (o3, Dec24/Apr25) → 17% (o4-mini, Apr25) → 38% Tiers1-3/19% Tier4 (Gemini 3 Pro, Nov25) → ~87-89% Tiers1-3 post-v2 correction (mid-2026) → 83% Tier4 v2 (Aug 2026).
3. **Extrapolation** — Code-execution analysis: naive fit (incl. o3 discontinuity) gives P(≥90%)≈20-52%; robust fit excluding that jump gives P≈0% (required pace ~4pp/month vs realized ~1.2-1.8pp/month). However this model is now stale — it predates the June 2026 v2 correction that already pushed scores to ~87-89%, making the extrapolation obsolete.
4. **Benchmark variant** — Market likely resolves on the main/Tiers1-3 leaderboard (already ~89%) rather than the harder Tier 4 v2 (~83%); description says "FrontierMath Exam" generically, primary source Epoch AI, so ambiguity remains on whether Tier 4 counts.
5. **Imminent releases** — GPT-5.6 family (Sol/Terra/Luna) and Claude Fable 5 already released and scoring 78-89%; no confirmed claim of ≥90% yet, but pace suggests imminent crossing.
6. **Prior benchmark saturation** — Not directly answered in research; general pattern (MATH, GPQA) saturates within 1-2 years of launch, consistent with Epoch researcher Greg Burnham's own view that FrontierMath will "probably saturate within the next two years — could be faster" (June 2026).
7. **Polymarket** — 89.5% (direct), up 7.5% over 30 days; no matching related Kalshi/Polymarket markets found (0 matches on keyword scans).
# Key facts
1. [arXiv] Nov 2024 launch SOTA <2%.
2. [Epoch AI Twitter] Apr 2025: o4-mini 17%, o3 10%.
3. [Epoch AI Twitter] Nov 2025: Gemini 3 Pro 38% (T1-3)/19% (T4).
4. [Epoch/GDELT] Jun 12 2026: v2 correction fixes errors in 42% of problems, raising scores broadly.
5. [llm-stats/benchlm] Jul 2026: GPT-5.6 Sol 89.0% (T1-3, main leaderboard).
6. [benchlm] Aug 2026: GPT-5.6 Sol 83% on corrected Tier 4 v2.
7. [Epoch AI] Jul 31 2026: New "Open Problems" expansion; only 3/50 solved.
8. [Polymarket direct] Current Polymarket YES = 89.5%.
# Cross-market signals
- Kalshi related: no direct FrontierMath match found; unrelated ERISA market only.
- Polymarket: 89.5% YES, rising 30-day trend, high volume for a niche market ($120k) — strong conviction.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Epoch researcher Greg Burnham (Jun 2026): expects saturation "within the next two years — could be faster."
- Digitalapplied blog: "today's top scores sit near 88%," implying saturation imminent.
- Caveat flagged by claude_news: much of the 87-89% data comes from third-party aggregators, not Epoch's own official page — real Epoch-certified SOTA may lag reported aggregator figures.
# Directional lean per outcome
- **Yes**: Strong — current aggregator SOTA already ~89% on main leaderboard, trend rising, Polymarket at 89.5%, only ~4 months remain to close and models are already at the threshold; pattern of step-jumps (o3, Gemini 3, GPT-5.5/5.6) suggests another leap plausible.
- **No**: Tier 4 (harder problems) still lags at ~83%; if resolution uses full/hardest variant or requires Epoch-official (not aggregator) confirmation, 90% may not be reached; stale trend-model calculations (excluding recent jump) implied near-0% probability, though outdated.
# Gaps / unknowns
- No kalshi_direct price was retrieved — anchor relies on Polymarket only.
- Ambiguity on whether "FrontierMath Exam" resolution uses Tiers1-3 only, full set, or includes Tier 4/Open Problems expansion.
- No official Epoch AI leaderboard snapshot confirming exact current SOTA (only third-party sites).
- Polymarket price discrepancy (89.5% vs. 83.5% in different tool outputs) unresolved.
# Calibration anchors
- Polymarket YES 89.5% (primary cross-market anchor, since Kalshi direct missing).
- Precedent: MATH/GPQA-style benchmarks saturated within ~1-2 years of launch; FrontierMath already showing similar trajectory 21 months post-launch.