← Back to scans

Will Anthropic have the #1 AI model at the end of September 2026?

0x77d9710dc8da78bc88f884327609f8546cc429374605bd117804bc4e9cdda2da · Science and Technology · 2026-08-14
75%
Agent
80%
Market Price
-5.5%
Edge
64%
Confidence
Volume: 19,440
Spread: 1.0c
Days to resolution: 47
Markets in event: 32
Final Rationale
Anthropic currently holds #1 on the exact resolution source (arena.ai Text Arena Overall, Style Control) via Claude Fable 5, and no confirmed rival flagship (OpenAI Astra/GPT-6, next-gen Gemini, further xAI) is dated before the 2026-09-30 snapshot, which supports Yes. The Polymarket proxy at 80.5% is the best available anchor but is thin (~$19K, 24 points) with a 34pp swing, so it deserves meaningful discounting rather than near-full weight. The red team is right that flip risk is underweighted: the top cluster sits within a few Elo points, leadership has shuffled every 1-3 months, Anthropic's own newest flagship (Opus 5) ranks only #6 on this specific board, and a re-baselining/methodology change (as occurred in July) is an additional tail risk independent of model quality. Balancing incumbency plus absence of scheduled competitor launches against noise-level margins and thin-market uncertainty, I settle modestly below both prior forecasts and well below the Polymarket anchor at 75% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 20$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related claude_news claude_news kalshi_related gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price for 'Anthropic #1 at end of September 2026', and how has it moved over the past 90 days?
  2. What are the current prices for the sibling markets in the same group (Google, OpenAI, xAI, Meta, DeepSeek, Other), and do they sum near 100% after de-vigging?
  3. Who currently occupies the #1 rank on arena.ai Text Arena (Overall, Style Control On), and by what margin over #2?
  4. How often has an Anthropic (Claude) model ever held the #1 LMArena text overall rank historically, and for how long?
  5. What frontier model releases from Anthropic (e.g., Claude 5 / Opus 4.5+) and competitors (Gemini 3.x/4, GPT-5.x, Grok 5) are expected between now and September 30, 2026?
  6. Does Anthropic reliably submit its newest models to LMArena, and how do Claude models typically rank under Style Control versus without it?
  7. How volatile is the #1 LMArena spot — how many distinct companies have held #1 over the trailing 12-18 months?
Planner reasoning
This is a Polymarket question about which company holds the #1 spot on the LMArena (arena.ai) Text Arena Overall leaderboard with Style Control on September 30, 2026. Key drivers are Anthropic's historical performance on LMArena (they have rarely held #1; Google Gemini and xAI/OpenAI dominate), the expected release cadence of frontier models before Sept 2026, and current market prices across all companies in the group. I'll anchor on Polymarket's own price, pull the sibling markets in the group (Google, OpenAI, xAI, etc.) for implied normalization, and research recent leaderboard standings.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will Anthropic have the #1 AI model at the end of September 2026?** - Current price (probability): 80.50% - 7-day price change: +13.00% - 30-day price change: +31.50% - Total volume: $19,440 (USD notional) - Price range: 46.50% - 84.00% - Data points: 24 days
polymarket_related OK 3.4s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword '#1 AI model end of September': 0 markets | keyword 'which company has #1 AI model': 0 markets | keyword 'best AI model': 0 markets | keyword 'Anthropic': 1 markets | keyword 'Gemini LMArena': 0 markets
claude_news OK 27.8s 12 Based on research into the LMArena (now "Arena," at arena.ai) Text Arena leaderboard: - **Current #1 (as of August 2026): Anthropic's Claude Fable 5**, holding the top spot on the Text Arena Overall leaderboard at roughly 1508-1525 Elo depending on the snapshot date. As of August 2026, Claude Fabl
claude_news OK 29.3s 10 Based on research as of mid-August 2026: - **Anthropic currently holds the #1 spot on the major aggregate leaderboard.** Anthropic released Claude Opus 5 on July 24, 2026, and on Artificial Analysis's leaderboard it is the top-ranked model overall, leading both the Intelligence Index at 63 and the
kalshi_related OK 3.3s 2 2 related markets / summaries. keyword 'AI model': ok | keyword 'LMArena': no matches | keyword 'best AI model': ok
gdelt_news OK 163.8s 20 GDELT: 20 articles across 3 queries (lookback=60d). 'LMArena leaderboard top model': 10 hits | 'Claude Opus LMArena rank': 10 hits | 'Gemini tops LMArena': error GDELT rate-limited after retries (429)
wikipedia OK 0.1s 4 Fetched 4 Wikipedia entries (0 missing pages).
code_execution OK 33.7s 0 ## Findings - **De-vig calculation**: Using illustrative sibling market prices (OpenAI 34¢, Google 30¢, Anthropic 19¢, xAI 10¢, Meta 4¢, Other 5¢), the raw sum is **1.02** (2% overround), implying a modest vig. - **Anthropic's fair market-implied probability** after normalizing to sum to 1: **18.6
3. Evidence Brief Sonnet · 7484 chars
# Current state Anthropic currently holds #1 on the arena.ai Text Arena (Overall) leaderboard via Claude Fable 5 (~1508-1525 Elo, style control), with Claude Opus 5 (released 2026-07-24) leading rival composite indices but ranking lower (#6) on pure LMArena human-preference voting. The market resolves on a snapshot check on 2026-09-30, so current standing is informative but not determinative — the top cluster (Anthropic, OpenAI, Google, xAI) is within a few Elo points, historically volatile. # Timeline of key events - 2026-05: Claude Opus 4.6 reported #1 on Text Arena at ~1418-1504 Elo (confirmed/reported, conflicting snapshots) — claude_news. - 2026-06-17: Cartesia claims #1 on Voice Arena (different board, not relevant) — gdelt. - 2026-06-24: Google's Gemini 3.5 Pro release reportedly slips to July (reported) — gdelt. - 2026-07-01/07-12: Arena leaderboard restored then re-baselined (methodology change causing temporary disruption) (confirmed) — claude_news. - 2026-07-19/20: Alibaba's Qwen3.8 claims #2 behind Claude Fable 5 (reported) — gdelt. - 2026-07-24: Anthropic launches Claude Opus 5 (confirmed) — gdelt/iclarified/arynews. - 2026-07-26: Kimi K3 open weights ship, leads Frontend Code Arena (reported) — claude_news. - 2026-07-31: GPT-5.6 family (Luna/Terra/Sol) joins official Text Arena (confirmed) — claude_news. - 2026-08-04-05: Fable 5 reported leading MirrorCode/other benchmarks; Qwen3.8-Max pricing news (reported) — gdelt. - 2026-08-12: xAI's Grok 4.6 launches, trails Anthropic/OpenAI on Artificial Analysis Index (confirmed) — claude_news. - 2026-08-14: Multiple trackers (sevenlab.ai, whatllm.org) confirm Claude Opus 5 leads composite rankings; Claude Fable 5 leads raw LMArena text board (reported) — claude_news. - No confirmed OpenAI "Astra"/GPT-6, Google next-gen Gemini, or additional xAI flagship dated before 2026-09-30 (reported absence) — claude_news. # Event Will Anthropic own the #1-ranked model on arena.ai Text Arena (Overall, Style Control On) at the 2026-09-30 12:00 PM ET check? # Outcomes to forecast Yes / No # Kalshi market anchor No direct kalshi_direct data was returned in this research pass (kalshi_related only surfaced unrelated markets, e.g., SI Swimsuit cover model). **Polymarket is the best available cross-market anchor**: current YES-equivalent price **80.5%**, up sharply from a 30-day low of 46.5% (+31.5% in 30 days, +13% in 7 days), on thin volume (~$19.4K total, 24 data points). This is a strong, recent bullish move toward Anthropic. # Sub-question answers 1. **Polymarket price/trend** — 80.5% currently; range 46.5%-84% over the sampled window; strong upward momentum in the last 7-30 days (polymarket_direct). 2. **Sibling market prices/de-vig** — Not available live; code_execution tool explicitly used **illustrative, fabricated placeholder prices** (OpenAI 33%, Google 29%, Anthropic 19%, xAI 10%, Meta 4%, Other 5%) — NOT real data, should be disregarded as evidence, only as a hypothetical framework. 3. **Current #1 rank holder** — Anthropic's Claude Fable 5 leads Text Arena Overall (Style Control) at ~1508-1550 Elo as of Aug 2026, per multiple secondary trackers; margin over #2 (Claude Opus 4.8/GPT-5.5 Pro/Gemini 3.1 Pro cluster) is only a few Elo points — statistically thin (claude_news). 4. **Historical Anthropic #1 tenure** — Anthropic has held #1 repeatedly through 2026 (Opus 4.6 → 4.7/4.8 → Fable 5 → possibly Opus 5), suggesting multi-month, non-continuous dominance in H1-H2 2026 (claude_news). Precise historical percentage not directly sourced; code_execution's "3 of 18 months" figure is illustrative, not verified. 5. **Upcoming frontier releases through Sept 2026** — OpenAI's next flagship ("Astra"/possible GPT-6) has no confirmed date; Google's Gemini 3.5/3.7 updates are incremental (Flash variants); xAI's Grok 4.6 already shipped (Aug 12) but trails top tier. No confirmed release expected to unseat Anthropic before close (claude_news). 6. **Anthropic's LMArena submission behavior / style control performance** — Anthropic reliably submits latest models (Fable 5, Opus 5 both listed); under Style Control, Claude Opus 4.6 hit 1550, confirming strong style-adjusted performance (claude_news). 7. **#1 volatility (12-18mo)** — Leadership has shuffled among Anthropic, OpenAI, Google, and briefly others; "four labs within Elo-noise of top" as of mid-2026, indicating high volatility historically (claude_news); Wikipedia/LMArena entry confirms multiple companies (OpenAI, Google DeepMind, Anthropic, plus DeepSeek pre-release testing) have used the platform, consistent with frequent leader changes. # Key facts (high-confidence, factual) 1. [claude_news] Anthropic's Claude Fable 5 leads Text Arena Overall (Style Control) at ~1508-1525 Elo as of August 2026. 2. [claude_news/gdelt] Anthropic launched Claude Opus 5 on 2026-07-24, currently #1 on several composite indices (Artificial Analysis) though #6 on pure LMArena text preference. 3. [claude_news] Top cluster (Anthropic, OpenAI GPT-5.6, Google Gemini 3.1 Pro) separated by only a few Elo points — within noise. 4. [claude_news] No confirmed OpenAI/Google/xAI flagship release before 2026-09-30 expected to clearly overtake Anthropic. 5. [Wikipedia] Anthropic valued at $965B (May 2026), reportedly planning IPO fall 2026 — signals strong momentum/resources. 6. [polymarket_direct] Polymarket YES price 80.5%, up from 46.5% low, reflecting a recent strong shift toward "Yes." # Cross-market signals - Kalshi related: no relevant direct data found this pass. - Polymarket: 80.5% YES, strong 30-day uptrend (+31.5pp), low volume (~$19.4K) — thin market, some noise risk. - Sportsbook implied: N/A. # Analyst opinions and speculation - claude_news synthesis: "well-positioned to still hold #1... though top tier is close enough a rival could flip it." - code_execution output explicitly flagged as illustrative/fabricated (not real prices/base rates) — should not be treated as evidence; its "~15-19%" estimate conflicts sharply with Polymarket's 80.5% and the qualitative evidence, likely because it used made-up sibling prices unrelated to real market state. # Directional lean per outcome - **Yes**: Anthropic currently occupies #1 on the exact resolution source (LMArena Style Control), has sustained #1 status across multiple 2026 releases, no confirmed competitor flagship threatens before close, Polymarket sentiment strongly bullish (80.5%, rising). - **No**: Historical volatility is high (leader changes every 1-3 months); Elo gaps are within noise (a few points), so a single close call or rival release (Gemini update, OpenAI Astra) could flip #1; 11-month-out uncertainty is inherently large. # Gaps / unknowns - No live Kalshi direct price captured this pass — true consensus anchor unknown beyond Polymarket proxy. - Sibling market (Google/OpenAI/xAI/etc.) real prices not obtained; de-vig analysis is fabricated/illustrative only. - Precise historical monthly leaderboard log (exact % of months Anthropic #1) not sourced from primary LMArena data. # Calibration anchors - Polymarket cross-market price: **80.5%** (rising sharply, thin volume caveat). - Current actual leaderboard state: Anthropic #1 (Claude Fable 5) as of Aug 2026 — direct, recent, and near the resolution mechanism itself. - Historical volatility precedent: #1 has changed hands among 3-4 labs within the trailing 12-18 months, suggesting meaningful (but not dominant) downside risk over the remaining ~6-7 weeks to close.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.62
Yes 78%
No 22%
Anthropic currently occupies #1 on the exact resolution source (arena.ai Text Arena Overall, Style Control) via Claude Fable 5 at ~1508-1525 Elo, and also leads composite indices with Claude Opus 5 (released 2026-07-24). No confirmed rival flagship (OpenAI 'Astra'/GPT-6, next-gen Gemini, further xAI models) is dated before the 2026-09-30 check; Grok 4.6 shipped Aug 12 and trails. The main downside is that the top cluster (GPT-5.6, Gemini 3.1 Pro) sits within a few Elo points — inside noise — and leadership has flipped among 3-4 labs repeatedly through 2026, so a re-baselining, a quiet model update, or a surprise release could unseat them in the remaining ~6-7 weeks. Polymarket's 80.5% YES (up from a 46.5% 30-day low, thin ~$19K volume) is the best available anchor; I shade slightly below it for thin-volume noise and residual flip risk, discarding the explicitly fabricated code_execution 15-19% figure.
gpt-5.5
0.66
Yes 76%
No 24%
Polymarket, the best available market anchor given no Kalshi-direct price, is at 80.5% YES and has moved sharply upward, which points toward Yes, though volume is thin. Anthropic currently holds #1 on the exact arena.ai Text Arena Overall Style Control source via Claude Fable 5, and its repeated 2026 leaderboard leadership plus reliable submissions point toward Yes. The main offset is that the top Anthropic/OpenAI/Google/xAI cluster is separated by only a few Elo points and historical #1 turnover has been frequent, which points toward No risk from noise or a surprise rival update. With no confirmed OpenAI, Google, or xAI flagship expected before the September 30 check, I stay bullish but shade below the 80.5% anchor for volatility and thin-market uncertainty.
5. Devil's Advocate Sonnet · red-team critique
1. **Underweighting flip frequency**: The brief notes "#1 has changed hands among 3-4 labs...repeatedly through 2026" and leaderboard flips happened roughly every 1-3 months (Opus 4.6 → Fable 5 → Opus 5 cluster shuffles just within May-Aug). Both forecasters land at 76-78% Yes, but a base rate of monthly-ish flips over a ~6-7 week remaining window arguably implies a No-probability closer to 30-35%, not 22-24%. 2. **Resolution-source mismatch not flagged**: Claude Opus 5 (Anthropic's flagship, launched 7/24) is explicitly #6 on pure LMArena Style Control despite leading composite indices — the actual #1 holder is the older Fable 5. Neither forecast discusses the risk that Anthropic's own model lineup could shift (e.g., a new competitor submission edges out Fable 5 specifically) even while "Anthropic" narratively still seems dominant. 3. **Convergence without independent signal**: Both forecasts (0.78 and 0.76) essentially just shade slightly below the same Polymarket anchor (80.5%) with nearly identical reasoning — neither brings an independent estimate (e.g., from historical flip-rate base rates or Elo-noise modeling) that could stress-test the anchor, so the two "independent" forecasts add little diversification value. 4. **Thin-volume caveat under-adjusted**: Polymarket volume is ~$19.4K across only 24 data points with a 34pp swing in 7-30 days — this is a strong signal of instability/manipulability, not settled consensus, yet both forecasts shade only ~2-5pp below it rather than pulling meaningfully toward a wider uncertainty band. 5. **Re-baselining/methodology risk omitted**: The brief notes a leaderboard "restored then re-baselined" event in July 2026 causing "temporary disruption" — neither forecast considers that another methodology change near the Sept 30 snapshot could arbitrarily alter rankings independent of actual model quality, which is a distinct tail risk from "rival releases a better model." 6. **No Kalshi-direct anchor acknowledged as a real gap in confidence**: Both forecasts note "no Kalshi direct data" but proceed as if Polymarket is a clean substitute; neither discounts confidence (e.g., via wider intervals or explicitly flagged uncertainty) for the fact that the actual resolution market's pricing is unknown and could differ from a thin cross-market proxy.
6. Aggregator Opus · final calibrated forecast
75% (64% confidence)
Yes: 75% No: 25%
Anthropic currently holds #1 on the exact resolution source (arena.ai Text Arena Overall, Style Control) via Claude Fable 5, and no confirmed rival flagship (OpenAI Astra/GPT-6, next-gen Gemini, further xAI) is dated before the 2026-09-30 snapshot, which supports Yes. The Polymarket proxy at 80.5% is the best available anchor but is thin (~$19K, 24 points) with a 34pp swing, so it deserves meaningful discounting rather than near-full weight. The red team is right that flip risk is underweighted: the top cluster sits within a few Elo points, leadership has shuffled every 1-3 months, Anthropic's own newest flagship (Opus 5) ranks only #6 on this specific board, and a re-baselining/methodology change (as occurred in July) is an additional tail risk independent of model quality. Balancing incumbency plus absence of scheduled competitor launches against noise-level margins and thin-market uncertainty, I settle modestly below both prior forecasts and well below the Polymarket anchor at 75% Yes.
Pipeline Timing
Total pipeline time: 264.5s
Per-tool research timings shown in the Research section above.