← Back to scans

AI model scores ≥ 90% on FrontierMath Benchmark before 2027?

0x4632e96b7010fa4c1474d876f178539f11a284761c08842db5dd2e737cf6341b · Companies · 2026-08-11
70%
Agent
83%
Market Price
-13.0%
Edge
low-medium
Confidence
Volume: 116,209
Spread: 2.0c
Days to resolution: 141
Markets in event: 1
Final Rationale
The brief indicates Polymarket's 83% YES is on this exact market, so it is a usable (if noisy) anchor, and it has been drifting down from 90% after the 88-89%→83% correction. The trajectory is genuinely steep — Tier 4 went from 39.6% (April 2026) to a verified 83% (Aug 2026) — and since Tier 4 is the hardest subset, Tiers 1-3 and hence the aggregate are likely far above the stale 52.4% April figure, which softens (though does not eliminate) the aggregate-vs-Tier-4 resolution ambiguity. Against that, the critique's strongest points stand: the only near-90% figure was walked back, no Epoch-certified ≥90% exists on any tier, the last 7-10 points of a benchmark are typically the hardest, FrontierMath was explicitly engineered to resist saturation, and Epoch's audit/verification cycle could push certification past Dec 31, 2026. I therefore settle modestly below the 83% anchor and roughly in line with (slightly below) both forecasters, at 70% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 21$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct article_search claude_news kalshi_related polymarket_related code_execution
Sub-questions (Fermi decomposition)
  1. What is the current polymarket market price for 0x4632e96b7010fa4c1474d876f178539f11a284761c08842db5dd2e737cf6341b?
  2. What recent news affects: AI model scores ≥ 90% on FrontierMath Benchmark before 2027??
  3. What do related prediction markets currently imply?
  4. What is the historical base rate for similar events?
Planner reasoning
planner fallback: broad sweep + category conditionals
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.7s 1 ## This Market's Polymarket Data **AI model scores ≥ 90% on FrontierMath Benchmark before 2027?** - Current price (probability): 83.00% - 7-day price change: -6.50% - 30-day price change: -5.00% - Total volume: $116,209 (USD notional) - Price range: 21.00% - 90.00% - Data points: 91 days
article_search FAILED 0.0s 0 tool unavailable (missing env var)
claude_news OK 38.1s 13 Here are key findings on FrontierMath benchmark progress: - **Benchmark overhauled mid-2026**: Epoch AI released FrontierMath v2 on June 12, 2026 after an audit found small but critical errors in 42% of the original problems , correcting 135 problems and removing 12, leaving 338 total. (digitalapp
kalshi_related OK 2.5s 1 1 related markets / summaries. keyword 'model': ok | keyword 'scores': no matches | keyword 'FrontierMath': no matches
polymarket_related OK 0.1s 3 Scanned 100 active Polymarket markets, kept 3 matches. keyword 'model': 1 markets | keyword 'scores': 1 markets | keyword 'FrontierMath': 0 markets | keyword 'Benchmark': 1 markets
code_execution OK 72.6s 0 **Key numerical findings** - **Anchor data points:** FrontierMath launched Nov 2024 with top models (GPT‑4o, Claude 3.5, Gemini 1.5) scoring **~2%**; OpenAI's o3 announcement (Dec 2024) claimed **~25%**, later verified/independent runs (Epoch AI, 2025) put leading models (o3, Gemini 2.5 Pro, Grok‑4
3. Evidence Brief Sonnet · 5963 chars
# Current state No model has a confirmed ≥90% score on FrontierMath as of the latest research (early August 2026). The best verified score on the hardest subset (Tier 4, v2) is 83% (GPT-5.6 Sol, OpenAI), after an initial 88-89% self-reported figure was corrected downward. The full/aggregate benchmark score (Tiers 1-4 combined) remains much lower (~40-52% for top models as of April 2026). Resolution hinges on EpochAI's official SOTA reporting, and it is unclear whether "the FrontierMath Exam" in the question means the aggregate score or a specific tier — a key ambiguity. # Timeline of key events - 2024-11: FrontierMath launches; top contemporary models (GPT-4o, Claude 3.5, Gemini 1.5) score ~2%. (confirmed, Galois/aiwiki) - 2024-12: OpenAI's o3 announced with ~25% score (initial claim). (confirmed, multiple sources) - 2025: Independent/Epoch-verified scores for leading models settle in 20-30% range on solvable tiers; Tier 4 (research-level, ~20% of problems) stays near 0-2% for all models. (reported, code_execution synthesis) - 2026-04: GPT-5.5 Pro reaches 52.4% on Tier 1-3, 39.6% on Tier 4 (official Epoch scoring per aiwiki/Wikipedia). (confirmed) - 2026-06-12: Epoch AI releases FrontierMath v2 after an audit found errors in 42% of original problems; 135 corrected, 12 removed, 338 problems remain. (confirmed, epoch.ai/digitalapplied.com) - 2026-08 (early): GPT-5.6 Sol leads FrontierMath v2 Tier 4 leaderboard; initial pre-release figure ~88-89% later corrected to 83% on public model re-run; GPT-5.6 Terra (68.3%), GPT-5.6 Luna (58.5%) follow. (reported/corrected, BenchLM.ai, X/Acer) - 2026-08: Commentators speculate a near-term Sol Pro variant "could score over 90% on FrontierMath Tier 4," but unconfirmed. (rumored, X/Twitter) # Event Will a SOTA AI model score ≥90% on the FrontierMath Exam (per EpochAI) before 2027 (close: 2026-12-31)? # Outcomes to forecast Yes / No # Kalshi market anchor No Kalshi-direct price was returned for this ticker in the research (only an unrelated "model" keyword match, a Sports Illustrated market). The Polymarket price (83% YES) is the closest available market anchor; treat with caution since it is not the Kalshi ticker itself. # Sub-question answers 1. **Current Polymarket price** — 83.00% YES, down 6.5% over 7 days and 5% over 30 days; range has been 21%-90% over 91 days, volume ~$116k (polymarket_direct). 2. **Recent news affecting this event** — FrontierMath v2 (June 2026) reset the leaderboard after error corrections; current Tier-4 leader GPT-5.6 Sol sits at a verified 83%, with an earlier 88-89% claim walked back. Full/aggregate benchmark scores remain far lower (~40-52%) as of April 2026 (claude_news). 3. **Related prediction markets** — No directly related Kalshi or Polymarket markets on FrontierMath found; unrelated matches (Bytedance "best AI model," SI cover model) are noise (kalshi_related, polymarket_related). 4. **Historical base rate** — Comparable hard benchmarks: MATH went from ~7% (2021) to ~90% (2024), taking ~3 years; GSM8K took ~2 years to go from 55% to 95%. FrontierMath was designed to resist saturation more than these, implying a longer timeline is more likely (code_execution). # Key facts (high-confidence, factual) 1. [claude_news/BenchLM] Best verified Tier-4 v2 score as of Aug 2026: 83% (GPT-5.6 Sol), corrected down from an initially reported 88-89%. 2. [aiwiki/Wikipedia] April 2026 aggregate (Tier 1-4) SOTA: 52.4% (GPT-5.5 Pro); Tier 4 alone: 39.6%. 3. [epoch.ai/digitalapplied] FrontierMath v2 (June 2026) corrected 135 problems, removed 12, from an audit revealing errors in 42% of original problems. 4. [polymarket_direct] Polymarket YES price 83%, trending down (-6.5% 7d, -5% 30d) from a peak of 90%. # Cross-market signals - Kalshi related: no meaningful matches found. - Polymarket: 83% YES on this exact market; declining trend suggests fading confidence after the 88-89%→83% correction. - Sportsbook implied: not applicable. # Analyst opinions and speculation - One X/Twitter commentator speculates a "Sol Pro" variant "could score over 90% on Tier 4" — unconfirmed rumor, not corroborated by benchmark trackers. - Epoch AI researcher expects FrontierMath to "probably saturate within the next two years – could be faster," implying eventual saturation but not necessarily by end-2026. - code_execution's logistic-curve extrapolation (using older/stale anchors) estimates 90% is reached around Jan-March 2027, i.e., just after the deadline, and assigns a probability of ~3-10% for "Yes" — but this model predates the newer 83% Tier-4 data point and may understate momentum. # Directional lean per outcome - **Yes**: Verified progress from 2% (Nov 2024) to 83% (Aug 2026) on the hardest tier is extremely fast; if this pace continues, crossing 90% on Tier 4 specifically by Dec 2026 is plausible within months. Speculative claims of imminent 90%+ variants add some support. - **No**: The 88-89% figure was walked back to 83%, showing self-reported/pre-release scores overstate progress; the full/aggregate benchmark (likely the more natural reading of "the FrontierMath Exam") remains far below 90% (~52%); historical analogous benchmarks took 2-3 years to go from similar levels to >90%; Tier 4's structural difficulty (0% for years) argues against a sudden near-complete leap. # Gaps / unknowns - No Kalshi-direct price was retrieved for this specific ticker — anchor relies on Polymarket only. - Ambiguity in resolution: does "FrontierMath Exam" mean aggregate score or Tier 4 subset? This materially changes probability (Tier 4 close to 90%; aggregate far from it). - No direct confirmation of any model hitting ≥90% on any tier as of the latest data (Aug 2026). # Calibration anchors - Polymarket YES price: 83% (declining trend, treat as noisy proxy anchor). - Historical analogs: MATH benchmark ~3 years from 7%→90%; GSM8K ~2 years from 55%→95%; FrontierMath explicitly designed to resist faster saturation than these.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
low
Yes 72%
No 28%
The closest market anchor (Polymarket, 83% YES but declining from a 90% peak) implies traders see this as likely, and the underlying trajectory supports that: FrontierMath v2 Tier 4 went from ~0-2% (2025) to a verified 83% (GPT-5.6 Sol, Aug 2026), and if the hardest tier is at 83%, the Tiers 1-3/main benchmark that Epoch headlines is almost certainly already near or above 90%, making the aggregate-vs-Tier-4 ambiguity less damaging than the brief's stale April 2026 figures suggest. Offsetting factors: the 88-89% claim was walked back to 83%, showing self-reported pre-release scores overstate progress; no ≥90% Epoch-verified SOTA is confirmed yet; and historical hard-benchmark analogs (MATH, GSM8K) took 2-3 years for the final climb, with FrontierMath explicitly engineered to resist saturation. I therefore shade modestly below the 83% Polymarket anchor to reflect the resolution ambiguity, the downward-correction precedent, and the declining price trend, while rejecting the stale logistic extrapolation's 3-10% as based on outdated anchors.
gpt-5.5
medium
Yes 68%
No 32%
The closest market anchor is Polymarket at 83% YES, but there is no Kalshi-direct price and the market has been declining after the 88-89% GPT-5.6 Sol figure was corrected to 83%. The current verified 83% on FrontierMath v2 Tier 4 points toward Yes because only a 7-point improvement is needed with several months remaining and progress has been extremely fast since 2024. However, the aggregate FrontierMath score was still far lower earlier in 2026, and the phrase 'FrontierMath Exam' may resolve on the full benchmark rather than the Tier 4 subset, which points materially toward No. Historical benchmark saturation timelines and the error-correction reset also argue against taking the 83% market price at face value, so I shade down from the anchor.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters treat the aggregate-vs-Tier-4 ambiguity as a modest haircut, but Forecast 1 goes further and simply asserts the aggregate "is almost certainly already near or above 90%" — this is speculative and directly contradicted by the brief's most recent hard data point (52.4% aggregate, April 2026, with no update since); this is an unjustified leap not supported by evidence. 2. Neither forecaster grapples with the fact that Epoch's own SOTA reporting (the resolution source) has not yet certified any ≥90% score on any tier, and that the only near-90% figure (88-89%) was explicitly walked back — this pattern of overstatement-then-correction should arguably justify more weight on "No" than either gives it, especially given only ~5 months remain to close. 3. Both anchor heavily on Polymarket's 83% despite the brief's explicit caution that it's "not the Kalshi ticker itself" and might resolve on different criteria/timing; neither discusses whether Polymarket's own resolution wording matches "per EpochAI" SOTA reporting, which could cause the proxy to be systematically biased. 4. The historical base-rate evidence (MATH: 3 years from 7%→90%; GSM8K: 2 years from 55%→95%; FrontierMath "designed to resist saturation") is acknowledged by both but effectively overridden by recency bias from the Aug 2026 83% figure — neither adequately reconciles why the current benchmark should buck its explicit design intent to resist rapid saturation. 5. Neither forecast addresses the possibility that Tier 4's "remaining problems" become disproportionately harder as scores climb (diminishing returns near ceiling), which is a common pattern in benchmark saturation (early gains are easy, final 10-20% is exponentially harder) — this argues for a lower probability than a naive linear/momentum extrapolation from 0%→83% would suggest. 6. Both forecasts converge closely (72% and 68%) despite using slightly different reasoning paths, but neither considers the resolution-timing risk: even if a model does cross 90% on some tier, EpochAI's official verification/reporting process (as shown by the June 2026 v2 audit and correction cycle) could introduce lag past the Dec 31, 2026 close, which independently pushes toward "No."
6. Aggregator Opus · final calibrated forecast
70% (low-medium confidence)
Yes: 70% No: 30%
The brief indicates Polymarket's 83% YES is on this exact market, so it is a usable (if noisy) anchor, and it has been drifting down from 90% after the 88-89%→83% correction. The trajectory is genuinely steep — Tier 4 went from 39.6% (April 2026) to a verified 83% (Aug 2026) — and since Tier 4 is the hardest subset, Tiers 1-3 and hence the aggregate are likely far above the stale 52.4% April figure, which softens (though does not eliminate) the aggregate-vs-Tier-4 resolution ambiguity. Against that, the critique's strongest points stand: the only near-90% figure was walked back, no Epoch-certified ≥90% exists on any tier, the last 7-10 points of a benchmark are typically the hardest, FrontierMath was explicitly engineered to resist saturation, and Epoch's audit/verification cycle could push certification past Dec 31, 2026. I therefore settle modestly below the 83% anchor and roughly in line with (slightly below) both forecasters, at 70% Yes.
Pipeline Timing
Total pipeline time: 183.1s
Per-tool research timings shown in the Research section above.