# Current state
Anthropic currently holds the #1, #2, and #3 positions on the Agent Arena "Models" leaderboard (Claude Fable 5, Claude Opus 5 Max, Claude Opus 5 High) as of early August 2026 [claude_news/cryptobriefing.com], a reversal from June 2026 when OpenAI's GPT-5.5 led. Resolution occurs on 2026-09-30 based on a leaderboard snapshot — a live, frequently-reshuffled ranking, not a fixed benchmark, so ~7 weeks of further model releases/rank churn remain before lock-in.
# Timeline of key events
- 2026-06 (reported): Agent Arena launches; #1 OpenAI GPT-5.5 (High), #2 Anthropic Claude-Opus-4.7, #3 Zai_org GLM-5.1, #4 Google Gemini-3.1-Pro [claude_news].
- 2026-06-29 (confirmed): TechCrunch reports Arena (LMArena rebrand) is now a $100M business [gdelt_news].
- 2026-07-09 (confirmed): Anthropic launches Claude Opus 5 [gdelt_news/anthropic.com].
- 2026-07-15 to 07-18 (reported): GPT-5.6 Sol narrows agentic gap with Claude Fable 5; Fable 5 becomes Anthropic's frontier public model [claude_news, gdelt_news].
- 2026-07-22 (confirmed): Google ships Gemini 3.6 Flash/Flash-Lite/Flash Cyber; Gemini 3.5 Pro still missing, missed June 2026 target [gdelt_news, claude_news].
- Early-mid Aug 2026 (reported): Anthropic's Claude Opus 5 Max climbs to #2 on Agent Arena, behind Fable 5; Opus 5 High takes #3 — Anthropic sweeps top 3 [cryptobriefing.com via claude_news].
- 2026-08-01 (reported): OpenAI names next major model "Astra" (informally "GPT-6"); unreleased [claude_news].
- 2026-08-07 (reported): Claude Opus 5 leads SWE-bench Verified (96%) and OSWorld 2.0 (70.6%) [benchlm.ai via claude_news].
# Event
Will Anthropic's model hold rank #1 on the Agent Arena "Models" leaderboard when checked 2026-09-30 12:00 PM ET?
# Outcomes to forecast
Yes / No (Anthropic best AI Agent by Sept 30, 2026)
# Kalshi market anchor
No kalshi_direct price was returned in research for this ticker — a gap. The only direct market-price data available is Polymarket: **73.5% YES** (current), down 2pts over 7 days but up 24pts over 30 days; range 44–79.5%; thin volume ($16.6K total, 21 data points). Treat this as a proxy anchor pending actual Kalshi YES price.
# Sub-question answers
1. **Who holds #1 now, by what margin?** Claude Fable 5 (Anthropic) is #1; Claude Opus 5 Max (#2) and Opus 5 High (#3) are also Anthropic — a top-3 sweep. GPT-5.6 Sol (OpenAI) is the nearest external competitor, described as "closing the gap" [cryptobriefing.com/claude_news].
2. **Historical churn of #1 spot?** #1 changed at least once in ~2 months: OpenAI GPT-5.5 led at June 2026 launch; Anthropic (Fable 5) took over by August 2026. No longer historical base rate available beyond this single observed flip; leaderboard appears volatile with monthly reshuffling as new model variants drop.
3. **Expected major releases through Sept 2026?** Google's Gemini 3.5 Pro/Gemini 4 delayed, no confirmed release before Sept 2026 (Nov/Dec estimates are analyst inference) [claude_news]. OpenAI's "Astra"/GPT-6 named Aug 1, 2026 but unreleased; Polymarket implies only ~71% chance of GPT-6 release by Sept 30 [claude_news]. xAI's Grok 5 (10T param) unlikely before Sept 2026; Grok 4.5/4.6 remains flagship [claude_news]. Anthropic has shown rapid cadence (Opus 4.5→5, Sonnet 5, Mythos, Fable) and is expected to continue iterating.
4. **Sibling Polymarket prices summing to ~1?** Not directly retrieved (polymarket_related returned 0 matches). A code_execution tool cited assumed/hypothetical sibling odds (OpenAI ~42%, Anthropic ~27%, Google ~22%, xAI ~6%) that conflict sharply with the actual Anthropic-market Polymarket price of 73.5% — these figures appear stale, hallucinated, or from an unrelated pull. **Flag: discount the code_execution de-vig analysis; the directly-observed 73.5% Polymarket price is the more reliable primary signal.**
5. **Claude vs Gemini/GPT on agentic benchmarks?** Claude Opus 5 leads SWE-bench Verified (96%) and OSWorld 2.0 (70.6%). Mixed on Terminal-Bench: GPT-5.6 Sol leads TB 65.9% vs Claude Fable 5's 62.9%; on TB 2.1, Claude Mythos 5 (88.0%) trails Kimi K3/GPT-5.6 Sol by ~0.5-0.8pts. On SWE-bench Pro (Scale standardized), GPT-5.4 leads; on vendor-aggregate active models, Opus 4.8 leads (69.2%). Gemini absent from top rankings across cited benchmarks — DeepMind lagging due to delays [claude_news x2].
6. **Is Agent Arena stable/comprehensive?** Yes, actively maintained by Arena.ai (formerly LMArena/Chatbot Arena, now a $100M business per TechCrunch), tracks OpenAI, Anthropic, Google, xAI (Grok), Zai_org (GLM), Kimi/Moonshot among others via real-world Agent Mode sessions — broad frontier coverage [wikipedia/LMArena, gdelt_news, claude_news].
# Key facts (high-confidence, factual)
1. [claude_news/cryptobriefing.com] Anthropic occupies Agent Arena ranks #1–3 as of early August 2026.
2. [claude_news] OpenAI led Agent Arena at its June 2026 launch (#1 GPT-5.5).
3. [claude_news] Gemini 3.5 Pro delayed past its June 2026 target; Gemini 4 has no confirmed date.
4. [claude_news] OpenAI's next flagship ("Astra") named but unreleased as of Aug 1, 2026.
5. [polymarket_direct] This exact market trades at 73.5% YES on Polymarket, up 24pts in 30 days.
6. [wikipedia] Anthropic valued at $965B (May 2026), Claude includes Opus/Sonnet/Mythos/Fable lines.
# Cross-market signals
- Kalshi related: "Will OpenAI or Anthropic IPO first?" favors Anthropic 83% — indicates market confidence in Anthropic's position generally, though unrelated to agent capability directly [kalshi_related].
- Polymarket (this market): 73.5% YES, rising trend (30d +24pts), but volatile (low $16.6K volume, wide 44–79.5% range) — thin liquidity reduces confidence in precision.
- Polymarket sibling markets: not found via search; code_execution's cited sibling odds are unverified/likely erroneous and contradict the direct 73.5% reading — do not treat as reliable.
# Analyst opinions and speculation
- cryptobriefing.com frames Anthropic as "competing against itself" at the top, implying a wide current lead.
- claude_news synthesis: bottom line assessment favors Yes, given Google/OpenAI/xAI major releases unlikely before Sept 30, 2026, though notes OpenAI's GPT-5.6 remains competitive on specific benchmarks (Terminal-Bench, SWE-bench Pro).
- Benchmark leadership called "highly volatile and fragmented" — no benchmark shows a stable, uncontested single leader across all agentic tasks [claude_news].
# Directional lean per outcome
- **Yes (Anthropic):** Current top-3 sweep on the exact resolution leaderboard; rapid Anthropic release cadence; competitors' major next-gen models (Gemini 4, GPT-6/Astra, Grok 5) unlikely before close; Polymarket trending up to 73.5%.
- **No (not Anthropic):** GPT-5.6 Sol already closing gap and leads on Terminal-Bench (original)/SWE-bench Pro standardized; leaderboard has flipped at least once in 2 months (OpenAI→Anthropic), showing real churn risk over remaining ~7 weeks; thin-volume Polymarket price may be noisy/unreliable; no Kalshi-direct anchor confirms consensus level.
# Gaps / unknowns
- No kalshi_direct YES price was returned — cannot confirm the actual Kalshi consensus, only proxied via Polymarket (73.5%).
- Sibling Polymarket markets (Google, OpenAI, xAI, Meta) not independently verified; code_execution figures likely unreliable/stale.
- No hard historical churn-rate data beyond one observed flip (June→August 2026); hazard-rate assumptions are speculative.
- Unclear whether new model releases (e.g., GPT-5.6 xHigh variants, Gemini updates) between now and Sept 30 could flip rank before close.
# Calibration anchors
- Polymarket YES price (proxy anchor): 73.5%, 30-day trend +24pts.
- Precedent: Agent Arena #1 already flipped once in ~2 months (OpenAI→Anthropic), suggesting non-trivial churn risk even over a short remaining window.
- Retention math (from code_execution, directionally useful): at monthly hazard 10–30% and ~7 weeks remaining, "current leader retains #1" probability roughly in the 65–85% range if Anthropic is assumed the true current leader.