# Current state
No model has a confirmed ≥90% score on FrontierMath as of the latest research (early August 2026). The best verified score on the hardest subset (Tier 4, v2) is 83% (GPT-5.6 Sol, OpenAI), after an initial 88-89% self-reported figure was corrected downward. The full/aggregate benchmark score (Tiers 1-4 combined) remains much lower (~40-52% for top models as of April 2026). Resolution hinges on EpochAI's official SOTA reporting, and it is unclear whether "the FrontierMath Exam" in the question means the aggregate score or a specific tier — a key ambiguity.
# Timeline of key events
- 2024-11: FrontierMath launches; top contemporary models (GPT-4o, Claude 3.5, Gemini 1.5) score ~2%. (confirmed, Galois/aiwiki)
- 2024-12: OpenAI's o3 announced with ~25% score (initial claim). (confirmed, multiple sources)
- 2025: Independent/Epoch-verified scores for leading models settle in 20-30% range on solvable tiers; Tier 4 (research-level, ~20% of problems) stays near 0-2% for all models. (reported, code_execution synthesis)
- 2026-04: GPT-5.5 Pro reaches 52.4% on Tier 1-3, 39.6% on Tier 4 (official Epoch scoring per aiwiki/Wikipedia). (confirmed)
- 2026-06-12: Epoch AI releases FrontierMath v2 after an audit found errors in 42% of original problems; 135 corrected, 12 removed, 338 problems remain. (confirmed, epoch.ai/digitalapplied.com)
- 2026-08 (early): GPT-5.6 Sol leads FrontierMath v2 Tier 4 leaderboard; initial pre-release figure ~88-89% later corrected to 83% on public model re-run; GPT-5.6 Terra (68.3%), GPT-5.6 Luna (58.5%) follow. (reported/corrected, BenchLM.ai, X/Acer)
- 2026-08: Commentators speculate a near-term Sol Pro variant "could score over 90% on FrontierMath Tier 4," but unconfirmed. (rumored, X/Twitter)
# Event
Will a SOTA AI model score ≥90% on the FrontierMath Exam (per EpochAI) before 2027 (close: 2026-12-31)?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct price was returned for this ticker in the research (only an unrelated "model" keyword match, a Sports Illustrated market). The Polymarket price (83% YES) is the closest available market anchor; treat with caution since it is not the Kalshi ticker itself.
# Sub-question answers
1. **Current Polymarket price** — 83.00% YES, down 6.5% over 7 days and 5% over 30 days; range has been 21%-90% over 91 days, volume ~$116k (polymarket_direct).
2. **Recent news affecting this event** — FrontierMath v2 (June 2026) reset the leaderboard after error corrections; current Tier-4 leader GPT-5.6 Sol sits at a verified 83%, with an earlier 88-89% claim walked back. Full/aggregate benchmark scores remain far lower (~40-52%) as of April 2026 (claude_news).
3. **Related prediction markets** — No directly related Kalshi or Polymarket markets on FrontierMath found; unrelated matches (Bytedance "best AI model," SI cover model) are noise (kalshi_related, polymarket_related).
4. **Historical base rate** — Comparable hard benchmarks: MATH went from ~7% (2021) to ~90% (2024), taking ~3 years; GSM8K took ~2 years to go from 55% to 95%. FrontierMath was designed to resist saturation more than these, implying a longer timeline is more likely (code_execution).
# Key facts (high-confidence, factual)
1. [claude_news/BenchLM] Best verified Tier-4 v2 score as of Aug 2026: 83% (GPT-5.6 Sol), corrected down from an initially reported 88-89%.
2. [aiwiki/Wikipedia] April 2026 aggregate (Tier 1-4) SOTA: 52.4% (GPT-5.5 Pro); Tier 4 alone: 39.6%.
3. [epoch.ai/digitalapplied] FrontierMath v2 (June 2026) corrected 135 problems, removed 12, from an audit revealing errors in 42% of original problems.
4. [polymarket_direct] Polymarket YES price 83%, trending down (-6.5% 7d, -5% 30d) from a peak of 90%.
# Cross-market signals
- Kalshi related: no meaningful matches found.
- Polymarket: 83% YES on this exact market; declining trend suggests fading confidence after the 88-89%→83% correction.
- Sportsbook implied: not applicable.
# Analyst opinions and speculation
- One X/Twitter commentator speculates a "Sol Pro" variant "could score over 90% on Tier 4" — unconfirmed rumor, not corroborated by benchmark trackers.
- Epoch AI researcher expects FrontierMath to "probably saturate within the next two years – could be faster," implying eventual saturation but not necessarily by end-2026.
- code_execution's logistic-curve extrapolation (using older/stale anchors) estimates 90% is reached around Jan-March 2027, i.e., just after the deadline, and assigns a probability of ~3-10% for "Yes" — but this model predates the newer 83% Tier-4 data point and may understate momentum.
# Directional lean per outcome
- **Yes**: Verified progress from 2% (Nov 2024) to 83% (Aug 2026) on the hardest tier is extremely fast; if this pace continues, crossing 90% on Tier 4 specifically by Dec 2026 is plausible within months. Speculative claims of imminent 90%+ variants add some support.
- **No**: The 88-89% figure was walked back to 83%, showing self-reported/pre-release scores overstate progress; the full/aggregate benchmark (likely the more natural reading of "the FrontierMath Exam") remains far below 90% (~52%); historical analogous benchmarks took 2-3 years to go from similar levels to >90%; Tier 4's structural difficulty (0% for years) argues against a sudden near-complete leap.
# Gaps / unknowns
- No Kalshi-direct price was retrieved for this specific ticker — anchor relies on Polymarket only.
- Ambiguity in resolution: does "FrontierMath Exam" mean aggregate score or Tier 4 subset? This materially changes probability (Tier 4 close to 90%; aggregate far from it).
- No direct confirmation of any model hitting ≥90% on any tier as of the latest data (Aug 2026).
# Calibration anchors
- Polymarket YES price: 83% (declining trend, treat as noisy proxy anchor).
- Historical analogs: MATH benchmark ~3 years from 7%→90%; GSM8K ~2 years from 55%→95%; FrontierMath explicitly designed to resist faster saturation than these.