← Back to scans

Will Anthropic have the #1 AI model at the end of September 2026?

0x77d9710dc8da78bc88f884327609f8546cc429374605bd117804bc4e9cdda2da · Science and Technology · 2026-09-01
87%
Agent
92%
Market Price
-4.5%
Edge
62%
Confidence
Volume: 26,295
Spread: 1.0c
Days to resolution: 29
Markets in event: 32
Final Rationale
The evidence strongly favors Yes: Anthropic's Fable 5 is reportedly #1 as of Aug 27 with multiple Anthropic models clustered near the top, only ~1 month remains, and the live Polymarket price (91.5%, rising) confirms the current state. The critique's strongest points — the illustrative 12.5% base rate and the ~13% de-vig — are both explicitly flagged as unverified/placeholder; the base rate measures unconditional occupancy, not the much higher conditional persistence given Anthropic is currently #1 with one month left. However, the critique is right that the 'depth' buffer argument is weak (a Gemini 3.5 Pro release could leapfrog all Anthropic models at once), that the anchor comes from a thin $26K market, and that the arena data is only secondarily sourced — justifying a slightly larger haircut than the consensus. I land at 0.87, modestly below both forecasters and the market anchor, reflecting the noise-level Elo margin and pending competitor releases before Sept 30.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 2$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 84% 87% 58%
2026-08-14 75% 80% 64%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. Which company's model currently holds rank #1 on the arena.ai Text Arena (Overall, Style Control On) leaderboard, and by what Elo margin over the next competitors?
  2. Where do Anthropic's latest Claude models currently rank on the Text Arena leaderboard with style control on, and what is their Elo gap to #1?
  3. What is Anthropic's historical base rate of holding #1 on LMArena/arena.ai, and how frequently has leadership changed hands in the past 12-24 months?
  4. Is Anthropic expected to release a new frontier model (e.g., next Claude generation) before September 30, 2026, and have past Claude releases debuted at or near #1 on the arena?
  5. What frontier releases are expected from Google (Gemini), OpenAI (GPT), and xAI (Grok) before end of September 2026 that could occupy #1?
  6. What probabilities do the sibling markets in this Polymarket group assign to Google, OpenAI, xAI, and others being #1, and do these sum consistently after de-vigging?
  7. Does Anthropic strategically prioritize arena performance, or do its models systematically underperform on arena rankings relative to benchmark/coding performance (style-control penalty effects)?
Planner reasoning
This is a Polymarket question about whether Anthropic will hold the #1 spot on the arena.ai Text Arena (Style Control On) leaderboard on Sept 30, 2026. The key drivers are the current leaderboard standings, Anthropic's historical arena performance (Claude models have rarely topped the arena vs Google/OpenAI), and expected model releases from all major labs before the check date. The market price is the primary anchor; cross-market prices for Google/OpenAI/xAI in the same group provide de-vigging context.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will Anthropic have the #1 AI model at the end of September 2026?** - Current price (probability): 91.50% - 7-day price change: +3.00% - 30-day price change: +12.00% - Total volume: $26,295 (USD notional) - Price range: 46.50% - 91.50% - Data points: 42 days
polymarket_related OK 2.8s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword '#1 AI model end of September': 0 markets | keyword 'best AI model September 2026': 0 markets | keyword 'Anthropic': 1 markets | keyword 'top AI model arena': 0 markets
kalshi_related OK 2.7s 1 1 related markets / summaries. keyword 'AI model arena': no matches | keyword 'best AI model': no matches | keyword 'Anthropic': ok
claude_news OK 36.5s 10 Here are the key findings on the arena.ai (formerly LMArena) leaderboard and frontier model landscape as of late August 2026: **Current #1 on Arena Text Leaderboard (Style Control) — Anthropic** - As of an August 27, 2026 snapshot, the Arena Overall Text leaderboard (7.9M+ votes, 395 models) is liv
gdelt_news OK 118.5s 20 GDELT: 20 articles across 4 queries (lookback=45d). 'Anthropic Claude LMArena leaderboard': 10 hits | 'Claude new model release 2026': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28claude+OR+new+OR+model+OR+release+OR+2026%2
wikipedia OK 0.1s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 42.0s 0 ## Findings **De-vig calculation (sibling markets)** - Illustrative raw YES prices used (placeholder — swap in live scraped prices when available): Google 48%, OpenAI 27%, xAI 12%, Anthropic 14%, Other 7% → sum = 108% (8-point vig/overlay). - Normalizing by dividing each by 1.08 gives de-vigged pro
3. Evidence Brief Sonnet · 6783 chars
# Current state As of the latest arena.ai snapshot (~2026-08-27, reported via secondary sources, not directly verified against arena.ai), Anthropic's **Claude Fable 5** sits at #1 on the Text Arena Overall (Style Control) leaderboard at ~1508-1525 Elo, but within a thin (~10-20 Elo) "noise band" of its own Opus 4.7/4.8 and rivals GPT-5.5 Pro / Gemini 3.1 Pro. The market itself (Polymarket, same question) prices Anthropic YES at **91.5%**, having risen sharply from a 46.5% low over the past 42 days — i.e., the tradeable market has already priced in Anthropic's current lead heavily, distinct from a separate illustrative de-vig calculation (see below) that conflicts with this. # Timeline of key events - 2025-11-24: Claude Opus 4.5 released; reported to "reclaim the coding crown from Gemini 3" (thenewstack.io, reported) - 2026-02 (approx): Claude Opus 4.6 released (claude_news synthesis, reported) - 2026-05-28: Claude Opus 4.8 released (Wikipedia/claude_news, reported) - 2026-07-01: Claude Fable "restoration" on arena boards (localaimaster.com, reported) - 2026-07-08: xAI ships Grok 4.5 flagship (reported) - 2026-07-09: GPT-5.6 "Sol" becomes default ChatGPT model (reported) - 2026-07-12: Arena score re-baseline event affecting Fable 5's listed Elo (reported) - 2026-07-24: Claude Opus 5 released; #1 on Artificial Analysis Intelligence/Agentic Index, top of 4 Arena boards incl. both coding boards (felloai.com, reported) - 2026-07-31: GPT-5.6 family (Luna, Terra, Sol) joins official Text Arena (reported) - 2026-08-12: xAI releases Grok 4.6 as value play, not leaderboard-focused; Gemini crosses 1B MAU (reported) - 2026-08-27: Arena snapshot shows Claude Fable 5 #1 at ~1508.6 Elo, GPT-5.6 Sol #14 (~1482.8) (felloai.com, reported) # Event Will Anthropic (via its top-ranked model) hold the #1 spot on arena.ai's Text Arena Overall (Style Control On) leaderboard when checked on 2026-09-30 12:00 PM ET? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct data was returned for this ticker; only Polymarket-direct data is available for this exact market. **Polymarket YES price: 91.5%**, up +3% (7d) and +12% (30d), off a 46.5% low over the observed 42-day window; volume is thin ($26.3K total). Treat this as the primary tradeable-market anchor in absence of Kalshi data. # Sub-question answers 1. **Current #1 & margin** — Claude Fable 5 (Anthropic) reportedly leads at ~1508-1525 Elo, with a tight cluster (Opus 4.8, GPT-5.5 Pro, Gemini 3.1 Pro, Opus 4.7) within ~10-20 Elo — effectively noise-level (claude_news/felloai.com/localaimaster.com, reported, not primary-verified). 2. **Claude's rank/gap** — Fable 5 is #1; Opus 4.8 and 4.7 sit just behind within the same noise band, meaning Anthropic occupies both #1 and several near-#1 slots (felloai.com). 3. **Base rate/churn** — Leadership has flip-flopped over the past year: Google's Gemini 3 briefly challenged Anthropic's coding lead before Opus 4.5 "reclaimed" it (Nov 2025); code_execution's illustrative Markov model assumes Anthropic held #1 only ~3/24 months historically (12.5% base rate) — this figure is explicitly labeled illustrative/unverified. 4. **New Anthropic releases expected** — Anthropic has already shipped Opus 5 (Jul 24) and Fable 5 (~Jul), both debuting at or near #1; rapid cadence (roughly every 2-3 months) suggests further releases plausible before Sept 30, 2026 close (claude_news/Wikipedia). 5. **Competing frontier releases** — OpenAI's GPT-5.6 family (Luna/Terra/Sol) is live since July 31; Google's Gemini 3.5 Pro remains undelivered as of late Aug 2026 (a lineup gap); xAI's Grok 4.5/4.6 positioned as price/value plays, not leaderboard contenders (medium.com/felloai.com). 6. **Sibling market de-vig** — code_execution used explicitly labeled "placeholder/illustrative" prices (Google 48%, OpenAI 27%, xAI 12%, Anthropic 14%, Other 7%), yielding de-vigged P(Anthropic)≈13% — this sharply contradicts the live Polymarket price (91.5%) for this exact question and should be treated as unreliable/stale, not real market data. 7. **Arena vs. benchmark focus** — Anthropic's Fable/Opus models top not just style-controlled text arena but also Arena's dedicated coding boards, suggesting genuine arena strength rather than a benchmark-only advantage undermined by style-control penalties (felloai.com). # Key facts (high-confidence, factual) 1. [Wikipedia] Anthropic's flagship lineup includes Haiku/Sonnet/Opus tiers plus a new "Fable"/"Mythos" line introduced in 2026. 2. [Polymarket] This market's live YES price is 91.5%, sharply up over 30 days. 3. [claude_news] Multiple secondary sources (felloai.com, localaimaster.com) report Claude Fable 5 at #1 on arena.ai Text Overall as of Aug 27, 2026. 4. [Kalshi_related] Separate Kalshi market "OpenAI or Anthropic IPO first" prices Anthropic at 90%, indicating market confidence in Anthropic's momentum broadly. # Cross-market signals - Kalshi related: Anthropic favored (90%) to IPO first vs OpenAI — general momentum signal, not arena-specific. - Polymarket (this market): 91.5% YES, strong uptrend. - Polymarket related: "Next Mythos-Class Model by Sep 1, 2026" priced 33.5% YES — suggests market skepticism on near-term next-gen release timing, though Fable/Opus 5 already released may satisfy earlier cadence. - Sportsbook implied: N/A. # Analyst opinions and speculation - felloai.com/localaimaster.com: Anthropic's lead is real but "within noise" (~10-20 Elo), meaning a single competitor release (Gemini 3.5 Pro, GPT-5.6 update) could flip #1 before Sept 30. - toolcenter.ai: Style Control explicitly reduces formatting bias, arguably favoring substance-driven scores — supports Anthropic's current genuine (not artifact-driven) lead. # Directional lean per outcome - **Yes (Anthropic #1)**: Supported by current arena snapshot (Fable 5 #1), rapid Anthropic release cadence, high Polymarket price (91.5%), and coding-board dominance. Opposing: extremely thin Elo margins, Gemini 3.5 Pro and GPT-5.6 updates still pending, historical churn. - **No**: Supported by narrow margin and unreleased competitor models (Gemini 3.5 Pro) that could overtake; opposed by strong current position and market pricing near-certainty. # Gaps / unknowns - No direct Kalshi YES price data retrieved for this ticker. - No primary arena.ai leaderboard scrape confirming Fable 5's rank/Elo — all evidence is secondary/blog-sourced. - Sibling-market de-vig figures are explicitly placeholder, not live. - Anthropic's exact historical #1-months count unverified. # Calibration anchors - Polymarket YES price (anchor): 91.5%, +12% over 30 days. - Precedent: Frontier #1 status has changed hands multiple times in 12 months (Gemini 3 vs. Opus 4.5, Nov 2025), suggesting non-trivial volatility despite current lead.
4. Forecasters Opus + GPT · in parallel
claude-fable-5
0.55
Yes 89%
No 11%
The Polymarket anchor prices YES at 91.5% with a strong uptrend, and the latest arena snapshot (Aug 27) shows Claude Fable 5 at #1 with Anthropic also holding multiple near-#1 slots (Opus 4.8/4.7), providing redundancy if any single model slips. Anthropic's lead spans both style-controlled text and coding boards, suggesting genuine strength rather than a formatting artifact. However, the ~10-20 Elo margin is within noise, Gemini 3.5 Pro remains unreleased and could land before Sept 30, and historical churn shows #1 status flips periodically — warranting a modest haircut off the market price. I land slightly below the anchor at 0.89 given only ~1 month remains and Anthropic's depth at the top of the board.
gpt-5.6-sol
0.68
Yes 88%
No 12%
The reported August 27 snapshot points toward Yes because Fable 5 is already #1 and multiple other Anthropic models occupy the near-leading cluster. Anthropic's recent release cadence and strength across both overall and coding leaderboards also favor Yes, while OpenAI's current GPT-5.6 models reportedly trail and Gemini 3.5 Pro remains unreleased. However, the 10–20 Elo margin is effectively noise-level, leadership has historically churned, and the snapshot is only secondarily sourced, leaving a meaningful chance that a competitor release or leaderboard movement flips the result. I therefore forecast slightly below the 91.5% Polymarket YES anchor, whose thin volume may overstate certainty.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts converge tightly around 0.88-0.89, almost exactly the Polymarket anchor minus a small haircut, but neither explains why the haircut should be ~2-3 points rather than larger given the brief explicitly flags a **12.5% base-rate figure** (Anthropic held #1 only ~3/24 months historically) — even if labeled illustrative, this is a stark reference-class contradiction neither forecaster grapples with quantitatively. 2. Neither forecast addresses the glaring inconsistency between the 91.5% Polymarket price and the sibling de-vig calculation implying ~13% P(Anthropic) — while the brief says treat the de-vig as unreliable, a rigorous forecast should explain *why* two market-derived signals diverge by ~78 points rather than simply dismissing one. 3. Both rely heavily on "Anthropic holds multiple near-#1 slots (Opus 4.8/4.7)" as a redundancy/buffer argument, but this reasoning is flawed: if the margin is genuinely noise-level (~10-20 Elo), a single strong Gemini 3.5 Pro or GPT-5.6 update could leapfrog *all* Anthropic models simultaneously, not just dislodge Fable 5 — the "depth" argument overstates protection against a real competitor jump. 4. Neither forecaster weighs the thin Polymarket volume ($26.3K) as a reason for skepticism about the 91.5% anchor's informational value — low liquidity markets can be driven by few participants and may not reflect well-calibrated aggregate belief, yet both treat it as a strong anchor. 5. The brief notes "Next Mythos-Class Model by Sep 1, 2026" is priced only 33.5% YES, suggesting market skepticism about Anthropic's near-term release cadence continuing — neither forecast incorporates this as a potential headwind against Anthropic maintaining its lead through fresh releases before the Sept 30 close. 6. Both forecasts land within 1 point of each other (0.89 vs 0.88) with reasoning that is nearly identical, suggesting insufficient independent analysis or diversity of models/methods — the near-identical outputs should raise concern about anchoring bias on the Polymarket price rather than genuine independent estimation from the Elo/base-rate evidence.
6. Aggregator Opus · final calibrated forecast
87% (62% confidence)
Yes: 87% No: 13%
The evidence strongly favors Yes: Anthropic's Fable 5 is reportedly #1 as of Aug 27 with multiple Anthropic models clustered near the top, only ~1 month remains, and the live Polymarket price (91.5%, rising) confirms the current state. The critique's strongest points — the illustrative 12.5% base rate and the ~13% de-vig — are both explicitly flagged as unverified/placeholder; the base rate measures unconditional occupancy, not the much higher conditional persistence given Anthropic is currently #1 with one month left. However, the critique is right that the 'depth' buffer argument is weak (a Gemini 3.5 Pro release could leapfrog all Anthropic models at once), that the anchor comes from a thin $26K market, and that the arena data is only secondarily sourced — justifying a slightly larger haircut than the consensus. I land at 0.87, modestly below both forecasters and the market anchor, reflecting the noise-level Elo margin and pending competitor releases before Sept 30.
Pipeline Timing
Total pipeline time: 225.6s
Per-tool research timings shown in the Research section above.