← Back to scans

Will Anthropic have the #1 AI model at the end of September 2026?

0x77d9710dc8da78bc88f884327609f8546cc429374605bd117804bc4e9cdda2da · Science and Technology · 2026-08-21
84%
Agent
87%
Market Price
-3.0%
Edge
58%
Confidence
Volume: 22,725
Spread: 2.0c
Days to resolution: 40
Markets in event: 32
Final Rationale
Anthropic holds the #1 slot now (Fable 5, ~1508–1525 Elo) and, critically, also appears to hold #2 (Opus 4.8) — a company-level resolution means intra-Anthropic succession or deprecation of Fable 5 mostly does NOT cost the win, which blunts the critique's 'self-cannibalization' concern (Opus 5's weak #6 text rank matters only if it displaces Anthropic's stronger text entrants, which is unlikely to remove Anthropic from #1 entirely). The verified Polymarket direct price of 87% is the only real anchor; I discount modestly for thin liquidity ($22.7k) and its wild 30-day swing (46.5%→87%), plus the genuinely narrow ~15–20 Elo cluster and two unreleased rival wildcards (OpenAI Astra with no date, Gemini 3.5 Pro reportedly months behind). Offsetting that, both rival flagships would have to ship AND leapfrog within six weeks, and Anthropic's release cadence is the fastest among majors, while its main historical loss of #1 was a regulatory suspension (now resolved with a safety classifier) rather than a competitive defeat. Arena re-baselining is a real methodological wildcard but is roughly symmetric in direction. Net: 0.84, a few points under the market to respect thin-market noise and the razor-thin recent margin without over-correcting against a durable, well-documented incumbency.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 13$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-14 75% 80% 64%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. Which company's model currently holds rank #1 on the arena.ai Text Arena (Overall) leaderboard with Style Control on, and by what score margin over #2?
  2. Where do Anthropic's best models (e.g., Claude Opus/Sonnet 4.5+) currently rank on the style-control text leaderboard, and what is the score gap to #1?
  3. How often has Anthropic held the #1 spot on LMArena text overall (style control) historically, and for how long at a time?
  4. What are the current Polymarket prices for the sibling markets (Google, OpenAI, xAI, Other) in this same event group, and what implied probability does Anthropic get after normalization?
  5. What major frontier model releases from Anthropic, Google, OpenAI, xAI, DeepSeek, and Meta are expected or rumored between now and end of September 2026?
  6. How volatile has the #1 position on LMArena been over the past 12 months (number of lead changes), which sets the base rate for the leader changing before Sept 30, 2026?
Planner reasoning
This is a Polymarket question about which company tops the LMArena (arena.ai) Text Overall Style-Control leaderboard on Sept 30, 2026. The key drivers are the current leaderboard #1, Anthropic's historical frequency of holding #1 (historically rare — Google/OpenAI have dominated text arena), and the expected model release pipeline through Q3 2026. Primary anchor is the Polymarket price plus the sibling markets for Google/OpenAI/xAI in the same event group, which should sum near 1 and let me back out an implied Anthropic probability.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will Anthropic have the #1 AI model at the end of September 2026?** - Current price (probability): 87.00% - 7-day price change: +7.50% - 30-day price change: +40.50% - Total volume: $22,725 (USD notional) - Price range: 46.50% - 87.00% - Data points: 31 days
polymarket_related OK 2.6s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword '#1 AI model end of September': 0 markets | keyword 'arena style control': 0 markets | keyword 'best AI model': 0 markets | keyword 'Anthropic': 1 markets | keyword 'Gemini OpenAI leaderboard': 0 markets
kalshi_related OK 2.5s 1 1 related markets / summaries. keyword 'best AI model': no matches | keyword 'LMArena': no matches | keyword 'Anthropic': ok
claude_news OK 25.5s 9 Based on multiple sources, here are the findings on LMArena/Arena.ai Text Arena Overall rankings (Style Control) as of August 2026: - **Anthropic currently holds #1**: As of August 2026, Claude Fable 5 (Anthropic) leads Arena's Overall Text leaderboard with Style Control at roughly 1508.6–1525 Elo.
claude_news OK 26.4s 15 ## Findings - **Anthropic (Claude) currently holds top rankings on multiple leaderboards as of mid-August 2026.** Anthropic released Claude Opus 5 on July 24, 2026, and on Artificial Analysis's leaderboard it is the top-ranked model overall, leading both the Intelligence Index at 63 and the Agenti
gdelt_news FAILED 240.0s 0 timeout after 240.0s
wikipedia OK 0.2s 4 Fetched 4 Wikipedia entries (0 missing pages).
code_execution OK 20.9s 0 **Findings (using illustrative sibling-market prices — replace with live Polymarket quotes for a precise read):** - **Raw sibling prices** (Google 45¢, OpenAI 27¢, Anthropic 15¢, xAI 6¢, Other 9¢) sum to **1.02**, implying a **~2% overround** typical of Polymarket's fee/liquidity spread. - **De-vig
3. Evidence Brief Sonnet · 7390 chars
# Current state The market resolves based on whichever company owns the #1-ranked model on the arena.ai Text Arena (Overall, Style Control) leaderboard on 2026-09-30 12:00 ET. As of the most recent research (mid-August 2026), Anthropic's Claude Fable 5 holds #1 (~1508–1525 Elo), but the top four models (Fable 5, Claude Opus 4.8, GPT-5.5 Pro, Gemini 3.1 Pro Preview) sit within ~15–20 Elo points of each other — a genuinely contestable margin with ~6 weeks remaining until close. # Timeline of key events - 2025-09 → 2026-08: Anthropic's Claude family (Opus 4.6→4.7→4.8→Fable 5) has occupied or closely contested #1 on the style-control text leaderboard for most of this ~12-month window (reported, claude_news synthesis). - 2026-02: LMArena reportedly showed Claude Opus 4.6 #1, GPT-5.2 #6 after OpenAI deprecated GPT-4o (reported, social-post cited by claude_news). - 2026-04-23: Text leaderboard kept Claude 4.6 variants on top; code leaderboard favored Claude Opus 4.7 variants (reported). - 2026-06-09: Claude Fable 5 launches, briefly tops board (~1525 Elo) (reported). - 2026-06-12: Fable 5 (and Claude Mythos 5) suspended worldwide under US export-control order targeting foreign-national access (confirmed via Wikipedia/claude_news). - 2026-07-01: Fable 5 access restored with enhanced safety classifier after 18-day suspension (confirmed/reported). - 2026-07-09: OpenAI ships GPT-5.6 (Sol/Terra/Luna tiers) as default ChatGPT model — 7th GPT-5.x release, no GPT-6 (reported). - 2026-07-12: Arena re-baselines Fable 5 score counting only post-restoration votes (reported). - 2026-07-24: Anthropic releases Claude Opus 5 — tops Artificial Analysis Intelligence/Agentic indices and Arena's coding/agentic/WebDev boards, but ranks only #6 on pure text arena (reported). - 2026-08-01: OpenAI teases next model "Astra" via solved problems — unreleased, no date/pricing/model card (reported). - 2026-08-07: LLM Stats snapshot shows GPT-5.6 Sol narrowly ahead of Claude Opus 5/Fable 5 on one composite index (reported, source-dependent). - 2026-08-12: xAI releases Grok 4.6, a post-training refresh (not new foundation model), trailing leaders (reported). - Ongoing (as of Aug 2026): Google's next flagship (Gemini 3.5 Pro) reported months behind schedule per Bloomberg; no confirmed launch date (reported). # Event Will Anthropic own the #1-ranked model on arena.ai Text Arena Overall (Style Control) as of Sept 30, 2026 12:00 ET? # Outcomes to forecast Yes / No # Kalshi market anchor No kalshi_direct price was returned for this ticker in the research (gap). Only a related Kalshi market was found: "Will OpenAI or Anthropic IPO first? — Anthropic" trading at 93% (up 8% over 30d, 1,047 avg daily contracts) — a proxy for general Anthropic momentum sentiment, not model-ranking specific. Treat this brief's cross-market anchor as Polymarket-based instead. # Sub-question answers 1. **Current #1 holder & margin** — Anthropic's Claude Fable 5 leads at ~1508–1525 Elo; margin to #2 (Claude Opus 4.8 / GPT-5.5 Pro cluster) is narrow, ~15–20 points (claude_news, Aug 2026 sources). 2. **Anthropic's best models' rank/gap** — Fable 5 is #1 on text; newer Claude Opus 5 (released 2026-07-24) ranks only #6 on pure text arena despite topping coding/agentic boards — Anthropic's frontier effort has bifurcated across model lines (claude_news). 3. **Historical #1 frequency** — Anthropic models have held or closely contested #1 for roughly 11 of the past 12 months (Sept 2025–Aug 2026), interrupted only by tight competition from Gemini 3.1 Pro/GPT-5.x and an 18-day forced export-control suspension (not a competitive loss) (claude_news synthesis). 4. **Polymarket sibling prices** — The code_execution tool's normalized sibling prices (Google 44%, OpenAI 27%, Anthropic 15%, xAI 6%, Other 9%) are explicitly labeled "illustrative," not live data — unreliable. The only verified live figure is Polymarket's direct quote on this Anthropic-outcome market: 87% YES. 5. **Expected releases through Sept 2026** — OpenAI's "Astra" announced but unreleased, no date (2026-08-01); Google's Gemini 3.5 Pro reported delayed/behind schedule, no date; xAI's Grok 4.6 (2026-08-12) is incremental only; Anthropic has shipped 4 Claude 5-family models in under two months (as of late July 2026), showing the fastest cadence among majors (claude_news). 6. **12-month volatility of #1** — Leadership has been relatively sticky (Anthropic-dominant), with brief close contests from Gemini/GPT-5.x but no confirmed multi-week loss of #1 to a rival, aside from Anthropic's own self-inflicted suspension (claude_news). Low structural lead-change base rate, but current Elo margins are thin enough that a single strong release (Astra, Gemini 3.5 Pro) could flip it. # Key facts (high-confidence, factual) 1. [polymarket_direct] This exact market trades at 87% YES, up from a 46.5% 30-day low (+40.5% 30d, +7.5% 7d), volume ~$22.7k (thin). 2. [Wikipedia] Anthropic is reportedly planning a fall 2026 IPO, valued at $965B (May 2026). 3. [Wikipedia] Anthropic faced a DoD dispute over surveillance/autonomous-weapons restrictions; phase-out blocked by injunction (2026-03-26). 4. [claude_news/Wikipedia] Fable 5 was suspended 2026-06-12 to 2026-07-01 under US export controls — a real, non-hypothetical Anthropic-specific tail risk. # Cross-market signals - Polymarket (this market): 87% YES, strong upward momentum. - Polymarket related: Anthropic $4.0T valuation by Dec 31 market only 10% YES — unrelated to model rank, shows valuation ≠ ranking confidence. - Kalshi related: Anthropic-vs-OpenAI IPO-first market at 93% — general Anthropic strength signal, not leaderboard-specific. - No credible normalized Polymarket sibling-market split obtained (only illustrative/synthetic figures returned). # Analyst opinions and speculation - Aggregator sites (felloai, localaimaster, buildfastwithai) consistently rank Anthropic (Fable 5/Opus 5) at or near #1 across multiple benchmarks as of Aug 2026, but flag that composite rankings are "source-dependent" and margins are within a few points. - Consensus view: frontier is a tight 4-way cluster (Anthropic, OpenAI, Google, xAI trailing), with Anthropic currently ahead but rivals' unreleased next-gen models (Astra, Gemini 3.5 Pro) are wildcards. # Directional lean per outcome - **Yes (Anthropic)**: Supported by ~11/12-month historical dominance, current #1 status, fastest release cadence, and 87% Polymarket price trending up. - **No (Other companies)**: Supported by extremely thin Elo margins (~15-20 pts), unreleased competitor flagships (OpenAI Astra, Gemini 3.5 Pro) that could leapfrog before Sept 30, and Anthropic's own regulatory fragility (precedent of forced suspension). # Gaps / unknowns - No direct Kalshi YES price captured for this ticker. - Reliability of specific 2026 model names/Elo scores ("Fable 5," "Opus 5," "Astra") not independently verified against official LMArena/Anthropic pages — treat as reported, not confirmed. - Sibling-market Polymarket data was synthetic/placeholder, not real. # Calibration anchors - Polymarket YES 87% (this market, real data) — primary anchor given Kalshi direct price missing. - Naive historical base rate ~11/12 ≈ 92% months Anthropic-led, roughly consistent with Polymarket price. - Tight Elo clustering (~15-20 pts) among top 4 injects meaningful uncertainty despite high base rate/market price.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.50
Yes 85%
No 15%
Anthropic currently holds #1 on the arena.ai Text Arena (Claude Fable 5, ~1508–1525 Elo) and has led or closely contested #1 for roughly 11 of the past 12 months, giving a high base rate for persistence. The primary live market anchor (Polymarket, this exact market) is 87% YES with strong upward momentum (+40.5% over 30d), consistent with the naive base rate. Offsetting risks are real but not decisive: the top four models cluster within ~15–20 Elo, OpenAI's teased 'Astra' and a delayed Gemini 3.5 Pro remain unreleased wildcards, and Anthropic showed regulatory fragility with the 18-day export-control suspension of Fable 5. I shade slightly below the 87% anchor because six weeks is enough for a single strong rival release to flip a thin margin and because the brief's model-level details are reported rather than independently verified.
gpt-5.5
0.66
Yes 84%
No 16%
With no Kalshi-direct price available, I anchor on the verified Polymarket direct price of 87% YES, while discounting slightly for thin volume and unresolved leaderboard/release uncertainty. The current #1-holder evidence points toward Yes because Anthropic's Claude Fable 5 is already leading the relevant Text Arena leaderboard, even though the 15–20 Elo margin is narrow. The historical-frequency and release-cadence evidence also point toward Yes because Anthropic has held or closely contested #1 for about 11 of the last 12 months and has shipped rapidly, while OpenAI Astra and Gemini 3.5 Pro remain unreleased or delayed. The main No case is that a single strong rival launch or another Anthropic-specific regulatory/access disruption could flip a very tight leaderboard before Sept. 30.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor almost entirely on the 87% Polymarket price with "thin volume ~$22.7k" and a 30-day range that swung from 46.5% to 87% — that volatility itself signals the market is unstable/low-information, yet neither forecaster discounts much for this beyond a token "slightly below" adjustment; a wide 30-day swing with thin liquidity should warrant a larger uncertainty haircut, not a ~2-3 point shave. 2. Neither forecast grapples with the key fact that Anthropic's *newest* flagship, Claude Opus 5 (released 2026-07-24), ranks only #6 on pure text arena despite topping coding/agentic boards — this shows Anthropic's frontier effort is bifurcating away from pure-text optimization, and if Fable 5 (the actual #1 holder) gets deprecated/succeeded by Opus-line before Sept 30, Anthropic could lose #1 to its own product-line transition, not just to a rival. This is a distinct "self-cannibalization" risk neither rationale addresses. 3. The historical base rate of "11 of 12 months" is used uncritically, but the brief notes the current lead is only ~15-20 Elo points among four clustered models — averaging over a 12-month period when the margin was likely much wider earlier inflates the base rate's relevance to the current razor-thin state; recent-margin-conditioned base rate would be lower than the naive 92%. 4. Neither forecast weights the re-baselining event (2026-07-12, Arena recounting only post-restoration votes for Fable 5) as a methodology risk — leaderboard recalculation policies could mechanically shift rankings independent of underlying model quality, an idiosyncratic risk specific to Arena's scoring methodology that neither rationale mentions. 5. Both converge to nearly identical probabilities (0.85 vs 0.84) via near-identical reasoning chains, showing no genuine independent triangulation — neither explores the "No" case seriously enough (e.g., quantifying probability that Astra or Gemini 3.5 Pro actually ships AND leapfrogs within 6 weeks, which the brief flags as a live wildcard given Google is "months behind schedule" but could still surprise). 6. Neither forecast addresses potential ambiguity in the resolution criterion itself — if Anthropic's #1 model at close is Opus 5 or a hybrid successor rather than Fable 5, and Arena's "Overall Style Control" ranking could plausibly shift between text-specialized vs. agentic-optimized models, there's unaddressed resolution-criteria risk beyond simple company-level win/loss.
6. Aggregator Opus · final calibrated forecast
84% (58% confidence)
Yes: 84% No: 16%
Anthropic holds the #1 slot now (Fable 5, ~1508–1525 Elo) and, critically, also appears to hold #2 (Opus 4.8) — a company-level resolution means intra-Anthropic succession or deprecation of Fable 5 mostly does NOT cost the win, which blunts the critique's 'self-cannibalization' concern (Opus 5's weak #6 text rank matters only if it displaces Anthropic's stronger text entrants, which is unlikely to remove Anthropic from #1 entirely). The verified Polymarket direct price of 87% is the only real anchor; I discount modestly for thin liquidity ($22.7k) and its wild 30-day swing (46.5%→87%), plus the genuinely narrow ~15–20 Elo cluster and two unreleased rival wildcards (OpenAI Astra with no date, Gemini 3.5 Pro reportedly months behind). Offsetting that, both rival flagships would have to ship AND leapfrog within six weeks, and Anthropic's release cadence is the fastest among majors, while its main historical loss of #1 was a regulatory suspension (now resolved with a safety classifier) rather than a competitive defeat. Arena re-baselining is a real methodological wildcard but is roughly symmetric in direction. Net: 0.84, a few points under the market to respect thin-market noise and the razor-thin recent margin without over-correcting against a durable, well-documented incumbency.
Pipeline Timing
Total pipeline time: 351.9s
Per-tool research timings shown in the Research section above.