# Current state
As of late August 2026, Anthropic's models (Claude Opus 5, Claude Fable 5, Claude Mythos 5) occupy the top 1-2 spots on the Agent Arena leaderboard (arena.ai/leaderboard/agent), per claude_news synthesis. Resolution occurs at a single snapshot (Sept 30, 2026, 12pm ET), so current leadership is informative but not determinative — no confirmed direct Kalshi YES price was returned in this research pull; Polymarket's parallel market prices Anthropic YES at 84.5%.
# Timeline of key events
- 2026-04 (reported): Claude Opus 4.7 leads SWE-bench Verified (87.6%); Claude Sonnet 4.5 leads GAIA; Anthropic sweeps top 6 GAIA spots. [claude_news/benchlm]
- 2026-05-28 (confirmed): Claude Opus 4.8 released. [claude_news]
- 2026-06-09 (confirmed): Claude Fable 5 released; tops newly-launched Agent Arena leaderboard by "widest margin ever" over Opus-4.8/GPT-5.5. [arena.ai X post]
- 2026-06 (confirmed): Arena.ai removes Claude Fable 5 following Anthropic announcement + US government directive to suspend access, despite it ranking #1 across Agent/Text/Code Arena. [arena.ai X post]
- 2026-06-30 (confirmed): Claude Sonnet 5 released. [claude_news]
- 2026-07-24 (confirmed): Claude Opus 5 launched; quickly scores 1500+ on Arena overall rankings, takes #1 on Text Arena; Anthropic reportedly holds majority of top-10 AA-Briefcase agentic slots. [marktechpost, artificialanalysis.ai]
- 2026-08-19 to 08-27 (reported): Agent Arena Pareto view shows Claude Opus 5 (High) #1, Claude Fable 5 (High) #2, Kimi K3 #3, GPT 5.5 #4, DeepSeek V4 Pro #5. [arena.ai/leaderboard/agent/pareto]
- 2026-08-27/28 (reported): SWE-bench Verified led by Claude Opus 5 (96%); but Terminal-Bench 2.0 led by GPT-5.6 Sol (91.9%) over Claude Mythos 5 (88.0%). [benchlm.ai]
# Event
Will Anthropic own the #1-ranked model on the Agent Arena leaderboard (arena.ai/leaderboard/agent, "Models" filter) at the Sept 30, 2026, 12pm ET snapshot?
# Outcomes to forecast
Yes (Anthropic #1), No (another company #1)
# Kalshi market anchor
No kalshi_direct price was returned in this research pull (tool output absent). Only a related Kalshi market was found: "Will OpenAI or Anthropic IPO first? — Anthropic" at 93% (unrelated to this question). **Treat Kalshi price as unknown/gap** — use Polymarket (84.5% YES for Anthropic) as the best available cross-market proxy anchor.
# Sub-question answers
1. **Current #1 and margin**: Claude Opus 5 (High) leads Agent Arena Pareto view at +12.47% net improvement, with Claude Fable 5 (High), also Anthropic, at +11.57% — a clear 1-2 sweep; Kimi K3 (Moonshot) is 3rd at +10.41%. [claude_news, arena.ai]
2. **Polymarket prices**: Only this exact market's Polymarket twin was found, pricing Anthropic YES at 84.5% (up from 44% low, +12pp in 30 days). No broader multi-company de-vigged breakdown was available; a code_execution tool attempted a hypothetical simulation (Anthropic ~27-35%) but explicitly used **illustrative, not real, prices** — disregard that figure.
3. **Turnover base rate**: Leadership has changed hands at least twice in ~3 months on this specific leaderboard (Fable 5 #1 in June → removed by government directive → Opus 5 reclaimed #1 in July), suggesting moderate-to-high turnover, but all changes have kept Anthropic on top except a brief regulatory-driven gap. [claude_news/arena.ai X]
4. **Upcoming releases**: No confirmed frontier releases named for Sept 2026 beyond current lineup (GPT-5.6 Sol/Terra/Luna, Gemini 3.1 Pro/Deep Think, Grok 4.5/4.6, Kimi K3, DeepSeek V4). OpenAI's GPT-5.6 Sol already leads Terminal-Bench 2.0; Gemini and Grok remain competitive on select benchmarks but not Agent Arena overall. [claude_news]
5. **Historical agentic lead**: Anthropic has led SWE-bench Verified continuously since ~April 2026 (Opus 4.7→Opus 5, 87.6%→96%) and dominates Agent Arena; but OpenAI leads Terminal-Bench 2.0 and a thin-sample Tau-bench is led by StepFun. Lead appears to be widening on SWE-bench/Agent Arena specifically, but is contested on Terminal-Bench. [benchlm.ai, marktechpost]
6. **Resolution-source risk**: Arena.ai actively maintains a changelog with frequent model additions (Opus 5 Max/High, GPT-5.6 variants, Kimi K3, Inkling), indicating active maintenance. However, precedent exists for models being pulled for regulatory reasons (Fable 5 removal in June 2026 following US government directive against Anthropic) — a real ambiguity/downside risk if it recurs near the Sept 30 snapshot. [arena.ai X, Wikipedia/Anthropic]
# Key facts (high-confidence, factual)
1. [claude_news/arena.ai] Anthropic holds #1 and #2 on Agent Arena Pareto leaderboard as of Aug 19-26, 2026.
2. [claude_news] Claude Opus 5 leads SWE-bench Verified at 96% (Aug 27, 2026).
3. [claude_news/benchlm] GPT-5.6 Sol leads Terminal-Bench 2.0 at 91.9%, ahead of Anthropic's Claude Mythos 5 (88.0%).
4. [Wikipedia] US government pressured DoD to phase out Anthropic products in Feb 2026 over autonomous-weapons/surveillance policy disputes; a federal injunction blocked this in March 2026. This same dynamic caused a temporary Arena removal of Fable 5 in June 2026.
5. [polymarket_direct] This exact market's Polymarket twin prices Anthropic YES at 84.5%, up 12pp over 30 days, on modest volume ($25.7k).
# Cross-market signals
- Kalshi related: No direct price found; unrelated Anthropic IPO market at 93% is not informative for this question.
- Polymarket: 84.5% YES for Anthropic (same-question twin market), trending up.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Medium reviewer (Aug 2026) frames Opus 5 as leading overall but notes GPT-5.6 Sol beats it on Terminal-Bench 2.1 and Agents' Last Exam — "close competitor" framing.
- claude_news synthesis concludes Anthropic "appears to have the strongest overall claim among trackers" but calls "best AI agent" benchmark-dependent/contested.
- code_execution's turnover-hazard model (using fabricated inputs) is unreliable and should not be weighted; flagged only as a methodology sketch.
# Directional lean per outcome
- **Yes (Anthropic)**: Strong — currently #1 and #2 on the exact resolution leaderboard, dominant on SWE-bench, rapid release cadence (4 major models since May), Polymarket at 84.5%.
- **No (other company)**: Some support — OpenAI leads Terminal-Bench 2.0; regulatory/government action against Anthropic caused one leaderboard removal already and remains an ongoing legal/political risk; one month of turnover already observed on this specific leaderboard.
# Gaps / unknowns
- No confirmed Kalshi direct YES price obtained in this pull.
- No genuine multi-outcome Polymarket "which company" breakdown (only binary Anthropic proxy).
- Unclear how Arena Arena's official "Models" filter (vs. Pareto view referenced in research) currently ranks companies — resolution uses "Models" tab specifically.
- Unknown whether regulatory action against Anthropic could recur before Sept 30, 2026 snapshot, creating resolution ambiguity.
# Calibration anchors
- Polymarket YES (twin market): 84.5%, uptrending.
- Current Agent Arena standing: Anthropic #1 and #2.
- Precedent: even under a June 2026 forced removal, Anthropic reclaimed #1 within ~6 weeks (Opus 5 launch July 24) — suggesting resilience of Anthropic's lead absent permanent bans.