# Event
Will Anthropic own the #1-ranked model on the Arena.ai "Agent Arena" Models leaderboard when checked Sept 30, 2026, 12:00 PM ET?
# Outcomes to forecast
- Yes (Anthropic ranked #1)
- No (any other company ranked #1)
# Kalshi market anchor
No kalshi_direct tool output was returned in this research pass (gap). The only direct price available is from Polymarket on the identical ticker: **78.5% YES**, up +7pp over 7 days and +26pp over 30 days (range 44%–79.5% over 32 days, ~$22.8k volume). Treat this as the best available consensus proxy until a live Kalshi quote is confirmed.
# Sub-question answers
1. **Who currently holds #1 on Agent Arena, by what margin?** As of the most recent verifiable snapshot (~Jul 21–Aug 19, 2026), Anthropic's Claude Fable 5 (High) leads with 12.72% net improvement vs. GPT-5.6 Sol (10.12%) and Claude Opus 4.8 (9.75%) — a ~2.6pp margin over the #2 non-Anthropic model [claude_news/manifold, arena.ai]. No confirmed data point exists for late September 2026.
2. **Turnover base rate over 6-12 months?** Since Agent Arena launched (~June 2026), Anthropic has held or shared #1 continuously: Claude Opus 4.8 debuted tied #1 with GPT-5.5 (June 2026), then Claude Fable 5 took sole #1 (Jul–Aug 2026) [x.com/arena]. No clean flip to a non-Anthropic sole leader has been observed in ~3 months of history — a small, Anthropic-favorable sample.
3. **Polymarket prices for each company / de-vig check?** Only Anthropic's own contract (78.5%) is directly observed; no companion Google/OpenAI/xAI markets were found (polymarket_related returned 0 matches). A code_execution "de-vig" exercise used **illustrative, not real, assumed prices** (Anthropic 46%) and should be disregarded as evidence — it is a hypothetical, not market data.
4. **Upcoming frontier releases before Sept 30, 2026?** GPT-6 (mid-Aug to mid-Sep 2026) and Claude Opus 5 successor cadence (Anthropic already shipped Opus 5 Jul 24, 2026) are both roadmapped for Q3; Gemini 4 seen as "more likely earlier" launch; Grok 5 (xAI) targeted Aug–Sep 2026 with high timing variance [digitalapplied.com, felloai.com].
5. **Agentic benchmark standing (SWE-bench, Terminal-Bench, OSWorld, tau-bench)?** As of Aug 2026, Anthropic (Opus 5/Mythos 5/Fable 5) leads SWE-bench Verified (96%), SWE-bench Pro (~80%), Terminal-Bench 3.0 (42.7% vs GPT-5.6 Sol 34.6%), and OSWorld 2.0/ARC-AGI-3/BrowseComp; OpenAI's GPT-5.6 Sol only leads on "classic coding" (DeepSWE) [benchlm.ai, codingfleet.com, datacamp.com]. Caveat: vendor-reported, methodology-inconsistent.
6. **Leaderboard idiosyncrasies?** Agent Arena uses live behavioral signals (user task-success labels, artifact downloads) and ranks by "net improvement," not raw Elo — a newer, less-standardized methodology than classic Chatbot Arena Elo [arena.ai/blog]. It's unclear how comprehensively all frontier labs (esp. xAI, Meta, DeepSeek) are represented; 51 models/1.9M sessions tracked as of Aug 19, 2026.
# Key facts (high-confidence, factual)
1. [claude_news/x.com] Anthropic models have held or shared #1 on Agent Arena continuously since its June 2026 launch through at least Aug 19, 2026.
2. [benchlm.ai, codingfleet.com] Anthropic leads most major agentic benchmarks (SWE-bench Verified/Pro, Terminal-Bench) as of Aug 2026.
3. [digitalapplied.com] GPT-6 and further Claude Opus releases are both roadmapped for Q3 2026, directly contesting the leaderboard near market close.
4. [Wikipedia] Anthropic valued at $965B (May 2026), planning IPO fall 2026 — signals strong momentum/resourcing but not directly resolution-relevant.
5. [polymarket_direct] This exact market prices Anthropic YES at 78.5%, trending up sharply (+26pp in 30 days).
# Cross-market signals
- Kalshi related: "Will OpenAI or Anthropic IPO first — Anthropic" trades 93% YES (unrelated to agent quality but signals strong market confidence in Anthropic generally).
- Polymarket (same ticker): 78.5% YES, rising trend.
- No sportsbook or other Polymarket sub-markets (per-company) found; polymarket_related scan returned zero matches.
# Analyst opinions and speculation
- toolcenter.ai: four labs (Anthropic, OpenAI, Google, xAI) are within "Elo-noise" of each other at the frontier — implies fragility of any lead.
- digitalapplied.com: GPT-6 and Opus successor launches "will set the agentic eval benchmark for the year," explicitly flagging Sept 2026 as a pivotal contested month.
- News coverage (GDELT) shows aggressive competitive activity from Google (Gemini 3.7 Flash, claims to beat Sonnet 5), Meta (Muse Code agent), DeepSeek (Code launch pending) — all could disrupt rankings before close.
# Directional lean per outcome
- **Yes (Anthropic)**: Supported by unbroken #1 streak since Arena launch, benchmark dominance across SWE-bench/Terminal-Bench/OSWorld, rising Polymarket price (78.5%, +26pp/30d), heavy investment/valuation momentum. Opposing: GPT-6 and Gemini 4 both plausibly launch before Sept 30, "Elo-noise" tightness among top 4 labs, single-source/short leaderboard history (only ~3 months of data).
- **No (other company)**: Supported by imminent GPT-6/Gemini 4/Grok 5 launches that could leapfrog Anthropic, historical precedent of leaderboard churn in adjacent arenas (e.g., LMArena text leaderboard already saw multiple lead changes), methodology opacity favoring surprise entrants (Kimi K3 nearly tied #1 in July). Opposing: no confirmed instance yet of a non-Anthropic model achieving sole #1 on this specific leaderboard.
# Gaps / unknowns
- No live Kalshi YES price captured in this research pass — must be reconciled with actual kalshi_direct data if available.
- No confirmed leaderboard snapshot closer than Aug 19, 2026; ~6-week gap to resolution date is significant given monthly release cadence.
- Polymarket per-company breakdown (Google/OpenAI/xAI shares) not found; the code_execution "de-vig" figures are fabricated/illustrative, not real market data — should not be treated as evidence.
- Unclear whether GPT-6, Gemini 4, or Grok 5 will actually ship and be indexed on Agent Arena before Sept 30 check time.
# Calibration anchors
- Polymarket YES on this exact ticker: 78.5% (rising).
- Precedent: Anthropic has held #1 for 100% of Agent Arena's ~3-month observed history — small-sample but directionally strong.
- Illustrative Markov turnover modeling (not real data) suggests market pricing implies an effective ~8-9%/month leadership flip rate — plausible but unverified.