← Back to scans

Will Anthropic have the best AI Agent at the end of September 2026?

0x432d3cd0bc52a9f24b76760c751a847b46efd803f1d1fcc60e86ce27ae207a3e · Science and Technology · 2026-08-29
82%
Agent
84%
Market Price
-2.0%
Edge
64%
Confidence
Volume: 25,718
Spread: 4.0c
Days to resolution: 33
Markets in event: 32
Final Rationale
Anthropic holds both #1 and #2 on the exact resolution leaderboard (Agent Arena) with a ~2pp margin over Kimi K3, leads SWE-bench Verified at 96%, and has shipped four frontier models since May 2026 — all strongly favoring Yes at a Sept 30 snapshot only ~5 weeks out. The same-question Polymarket twin at 84.5% (uptrending) is the best available anchor absent a Kalshi direct price, but its thin $25.7k volume warrants a modest discount rather than trading above it. The devil's advocate correctly notes three under-weighted downside channels: (a) the resolution uses the 'Models' filter while the evidence cites the Pareto view, creating criteria-mismatch risk; (b) live regulatory/injunction dynamics already forced one Arena removal of a #1-ranked Anthropic model in June 2026 and could recur; and (c) a single-snapshot mechanism is more vulnerable to a surprise release (OpenAI already leads Terminal-Bench 2.0 and Agents' Last Exam) given the observed ~monthly frontier cadence. These push me slightly below both forecasts and marginally below Polymarket, to 82% Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 5$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-21 76% 78% 56%
2026-08-10 75% 74% 38%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related claude_news claude_news gdelt_news kalshi_related code_execution wikipedia
Sub-questions (Fermi decomposition)
  1. Which company's model currently occupies rank 1 on the Agent Arena leaderboard (arena.ai/leaderboard/agent, Models filter), and by what margin?
  2. What are the current Polymarket prices for each company in the 'which company has the best AI agent end of September 2026' group, and what does the de-vigged implied probability for Anthropic look like?
  3. How frequently has the #1 spot on the Agent Arena / LMArena agentic leaderboard changed hands over the past 6-12 months (base rate of leader turnover)?
  4. What major frontier model releases (Claude, Gemini, GPT, Grok, Qwen/DeepSeek) are expected or rumored between now and September 2026, and how have their agentic/tool-use benchmark scores trended?
  5. Has Anthropic historically led agentic coding/tool-use benchmarks (SWE-bench, Terminal-Bench, tau-bench) relative to Google and OpenAI, and is that lead widening or narrowing?
  6. Is arena.ai's Agent Arena leaderboard actively maintained/updated, and are there risks of resolution-source ambiguity (ties, delisted models, site downtime)?
Planner reasoning
This is a Polymarket question resolving on the Agent Arena leaderboard (arena.ai) top rank on Sept 30, 2026, so the primary anchor is the Polymarket price plus the sibling markets for Google/OpenAI/xAI etc., which should sum to ~1 after de-vigging. The key empirical inputs are the current leaderboard leader, Anthropic's historical hold on agentic benchmarks, and the expected model release cadence (Claude next-gen vs Gemini/GPT) over the ~11 months to resolution.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will Anthropic have the best AI Agent at the end of September 2026?** - Current price (probability): 84.50% - 7-day price change: +1.00% - 30-day price change: +12.00% - Total volume: $25,718 (USD notional) - Price range: 44.00% - 84.50% - Data points: 40 days
polymarket_related OK 3.4s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'best AI agent': 0 markets | keyword 'Agent Arena': 0 markets | keyword 'best AI model end of': 0 markets | keyword 'Anthropic': 0 markets | keyword 'Google Gemini best model': 0 markets
claude_news OK 28.7s 15 Based on my research, here are the current findings on the Agent Arena leaderboard (arena.ai/leaderboard/agent): - **Current #1 (as of Aug 19-26, 2026)**: Claude Opus 5 (High) by Anthropic leads the Pareto-optimality view with +12.47% net improvement at $2.48/task , followed by Claude Fable 5 (Hig
claude_news OK 34.7s 14 ## Findings on AI Agent Leadership Heading into September 2026 - **Anthropic has released rapidly through 2026**: Claude Opus 4.5 (Nov 2025) → Opus 4.6 → Opus 4.7 → Opus 4.8 (May 28, 2026) → Sonnet 5 (June 30) → Claude Fable 5 (June 9) → Claude Opus 5 (July 24, 2026), with "Claude Mythos" as a prev
gdelt_news OK 3.5s 0 GDELT: 0 articles across 3 queries (lookback=60d). 'Agent Arena leaderboard arena.ai': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28agent+OR+arena+OR+leaderboard+OR+arena%29+sourcelang%3Aenglish&mode=ArtList&format=json&sta
kalshi_related OK 3.2s 1 1 related markets / summaries. keyword 'best AI model': no matches | keyword 'AI agent': no matches | keyword 'Anthropic': ok
code_execution OK 28.5s 0 **Note:** No live Polymarket price feed was supplied in the prompt, so the figures below use a representative illustrative price set (in cents/probability) for a 6-way "best AI agent, end of Sep 2026" market — Anthropic 28¢, OpenAI 48¢, Google DeepMind 15¢, xAI 6¢, Meta 3¢, Other/Field 2¢ — to demon
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
3. Evidence Brief Sonnet · 7267 chars
# Current state As of late August 2026, Anthropic's models (Claude Opus 5, Claude Fable 5, Claude Mythos 5) occupy the top 1-2 spots on the Agent Arena leaderboard (arena.ai/leaderboard/agent), per claude_news synthesis. Resolution occurs at a single snapshot (Sept 30, 2026, 12pm ET), so current leadership is informative but not determinative — no confirmed direct Kalshi YES price was returned in this research pull; Polymarket's parallel market prices Anthropic YES at 84.5%. # Timeline of key events - 2026-04 (reported): Claude Opus 4.7 leads SWE-bench Verified (87.6%); Claude Sonnet 4.5 leads GAIA; Anthropic sweeps top 6 GAIA spots. [claude_news/benchlm] - 2026-05-28 (confirmed): Claude Opus 4.8 released. [claude_news] - 2026-06-09 (confirmed): Claude Fable 5 released; tops newly-launched Agent Arena leaderboard by "widest margin ever" over Opus-4.8/GPT-5.5. [arena.ai X post] - 2026-06 (confirmed): Arena.ai removes Claude Fable 5 following Anthropic announcement + US government directive to suspend access, despite it ranking #1 across Agent/Text/Code Arena. [arena.ai X post] - 2026-06-30 (confirmed): Claude Sonnet 5 released. [claude_news] - 2026-07-24 (confirmed): Claude Opus 5 launched; quickly scores 1500+ on Arena overall rankings, takes #1 on Text Arena; Anthropic reportedly holds majority of top-10 AA-Briefcase agentic slots. [marktechpost, artificialanalysis.ai] - 2026-08-19 to 08-27 (reported): Agent Arena Pareto view shows Claude Opus 5 (High) #1, Claude Fable 5 (High) #2, Kimi K3 #3, GPT 5.5 #4, DeepSeek V4 Pro #5. [arena.ai/leaderboard/agent/pareto] - 2026-08-27/28 (reported): SWE-bench Verified led by Claude Opus 5 (96%); but Terminal-Bench 2.0 led by GPT-5.6 Sol (91.9%) over Claude Mythos 5 (88.0%). [benchlm.ai] # Event Will Anthropic own the #1-ranked model on the Agent Arena leaderboard (arena.ai/leaderboard/agent, "Models" filter) at the Sept 30, 2026, 12pm ET snapshot? # Outcomes to forecast Yes (Anthropic #1), No (another company #1) # Kalshi market anchor No kalshi_direct price was returned in this research pull (tool output absent). Only a related Kalshi market was found: "Will OpenAI or Anthropic IPO first? — Anthropic" at 93% (unrelated to this question). **Treat Kalshi price as unknown/gap** — use Polymarket (84.5% YES for Anthropic) as the best available cross-market proxy anchor. # Sub-question answers 1. **Current #1 and margin**: Claude Opus 5 (High) leads Agent Arena Pareto view at +12.47% net improvement, with Claude Fable 5 (High), also Anthropic, at +11.57% — a clear 1-2 sweep; Kimi K3 (Moonshot) is 3rd at +10.41%. [claude_news, arena.ai] 2. **Polymarket prices**: Only this exact market's Polymarket twin was found, pricing Anthropic YES at 84.5% (up from 44% low, +12pp in 30 days). No broader multi-company de-vigged breakdown was available; a code_execution tool attempted a hypothetical simulation (Anthropic ~27-35%) but explicitly used **illustrative, not real, prices** — disregard that figure. 3. **Turnover base rate**: Leadership has changed hands at least twice in ~3 months on this specific leaderboard (Fable 5 #1 in June → removed by government directive → Opus 5 reclaimed #1 in July), suggesting moderate-to-high turnover, but all changes have kept Anthropic on top except a brief regulatory-driven gap. [claude_news/arena.ai X] 4. **Upcoming releases**: No confirmed frontier releases named for Sept 2026 beyond current lineup (GPT-5.6 Sol/Terra/Luna, Gemini 3.1 Pro/Deep Think, Grok 4.5/4.6, Kimi K3, DeepSeek V4). OpenAI's GPT-5.6 Sol already leads Terminal-Bench 2.0; Gemini and Grok remain competitive on select benchmarks but not Agent Arena overall. [claude_news] 5. **Historical agentic lead**: Anthropic has led SWE-bench Verified continuously since ~April 2026 (Opus 4.7→Opus 5, 87.6%→96%) and dominates Agent Arena; but OpenAI leads Terminal-Bench 2.0 and a thin-sample Tau-bench is led by StepFun. Lead appears to be widening on SWE-bench/Agent Arena specifically, but is contested on Terminal-Bench. [benchlm.ai, marktechpost] 6. **Resolution-source risk**: Arena.ai actively maintains a changelog with frequent model additions (Opus 5 Max/High, GPT-5.6 variants, Kimi K3, Inkling), indicating active maintenance. However, precedent exists for models being pulled for regulatory reasons (Fable 5 removal in June 2026 following US government directive against Anthropic) — a real ambiguity/downside risk if it recurs near the Sept 30 snapshot. [arena.ai X, Wikipedia/Anthropic] # Key facts (high-confidence, factual) 1. [claude_news/arena.ai] Anthropic holds #1 and #2 on Agent Arena Pareto leaderboard as of Aug 19-26, 2026. 2. [claude_news] Claude Opus 5 leads SWE-bench Verified at 96% (Aug 27, 2026). 3. [claude_news/benchlm] GPT-5.6 Sol leads Terminal-Bench 2.0 at 91.9%, ahead of Anthropic's Claude Mythos 5 (88.0%). 4. [Wikipedia] US government pressured DoD to phase out Anthropic products in Feb 2026 over autonomous-weapons/surveillance policy disputes; a federal injunction blocked this in March 2026. This same dynamic caused a temporary Arena removal of Fable 5 in June 2026. 5. [polymarket_direct] This exact market's Polymarket twin prices Anthropic YES at 84.5%, up 12pp over 30 days, on modest volume ($25.7k). # Cross-market signals - Kalshi related: No direct price found; unrelated Anthropic IPO market at 93% is not informative for this question. - Polymarket: 84.5% YES for Anthropic (same-question twin market), trending up. - Sportsbook implied: N/A. # Analyst opinions and speculation - Medium reviewer (Aug 2026) frames Opus 5 as leading overall but notes GPT-5.6 Sol beats it on Terminal-Bench 2.1 and Agents' Last Exam — "close competitor" framing. - claude_news synthesis concludes Anthropic "appears to have the strongest overall claim among trackers" but calls "best AI agent" benchmark-dependent/contested. - code_execution's turnover-hazard model (using fabricated inputs) is unreliable and should not be weighted; flagged only as a methodology sketch. # Directional lean per outcome - **Yes (Anthropic)**: Strong — currently #1 and #2 on the exact resolution leaderboard, dominant on SWE-bench, rapid release cadence (4 major models since May), Polymarket at 84.5%. - **No (other company)**: Some support — OpenAI leads Terminal-Bench 2.0; regulatory/government action against Anthropic caused one leaderboard removal already and remains an ongoing legal/political risk; one month of turnover already observed on this specific leaderboard. # Gaps / unknowns - No confirmed Kalshi direct YES price obtained in this pull. - No genuine multi-outcome Polymarket "which company" breakdown (only binary Anthropic proxy). - Unclear how Arena Arena's official "Models" filter (vs. Pareto view referenced in research) currently ranks companies — resolution uses "Models" tab specifically. - Unknown whether regulatory action against Anthropic could recur before Sept 30, 2026 snapshot, creating resolution ambiguity. # Calibration anchors - Polymarket YES (twin market): 84.5%, uptrending. - Current Agent Arena standing: Anthropic #1 and #2. - Precedent: even under a June 2026 forced removal, Anthropic reclaimed #1 within ~6 weeks (Opus 5 launch July 24) — suggesting resilience of Anthropic's lead absent permanent bans.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.60
Yes 86%
No 14%
Anthropic currently holds both #1 and #2 on the exact resolution leaderboard (Agent Arena), leads SWE-bench Verified at 96%, and has shipped four major models since May 2026, showing rapid cadence that makes displacement in ~5 weeks less likely. The Polymarket twin market prices Anthropic YES at 84.5% and is trending up (+12pp in 30 days), which serves as the best available anchor absent a Kalshi direct price. Downside risks are real but bounded: GPT-5.6 Sol leads Terminal-Bench 2.0, and a June 2026 US-government-directive-driven removal of Claude Fable 5 from Arena shows a nonzero regulatory tail risk; however, Anthropic reclaimed #1 within six weeks after that event. No confirmed frontier competitor release is scheduled for September 2026 that specifically threatens Agent Arena leadership. I settle slightly above the Polymarket price given the depth of Anthropic's 1-2 sweep and margin over Kimi K3.
gpt-5.6-sol
0.68
Yes 83%
No 17%
Anthropic currently holds the top two Agent Arena Pareto positions, with Opus 5 leading Kimi K3 by roughly two percentage points in net improvement, pointing strongly toward Yes. Its sustained SWE-bench leadership and lack of confirmed rival frontier releases before the snapshot further support persistence through September, although OpenAI's Terminal-Bench lead shows meaningful competitive risk. The leaderboard's turnover and the prior regulatory-driven removal of Fable 5 create downside, while uncertainty over whether the resolution's Models filter matches the cited Pareto ranking warrants an additional discount. I therefore forecast 83% Yes, slightly below the same-question Polymarket proxy anchor of 84.5%.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor heavily on Polymarket (84.5%) and land within a narrow 83-86% band, but the brief explicitly flags this as a thin-volume ($25.7k) twin market that may not reflect true probability — neither forecaster discounts for potential mispricing or manipulation risk on such low liquidity, nor do they address why Kalshi's own price (the actual resolution venue) is unknown, which is a notable anchoring gap given the instruction to prioritize Kalshi YES price. 2. Neither forecast adequately weights the "Models" filter vs. "Pareto view" ambiguity flagged in the brief ("Unclear how Arena Arena's official 'Models' filter... currently ranks companies — resolution uses 'Models' tab specifically") — Forecast 2 mentions it only as a passing "additional discount" without quantifying how much this resolution-criteria mismatch should matter, and Forecast 1 ignores it entirely. 3. Both forecasts treat the regulatory/government risk (DoD phase-out pressure, the June 2026 forced removal of Fable 5) as a bounded, already-resolved tail risk based on one past resolution, but the brief notes this is an "ongoing legal/political risk" with a live federal injunction dynamic — a recurrence within the 5-week window before the specific Sept 30 snapshot could cause an outright removal or ambiguous resolution, which deserves more explicit probability mass than a passing mention. 4. Neither forecaster stress-tests the scenario where GPT-5.6 Sol or another model overtakes Opus 5 specifically on the Agent Arena "Models" leaderboard given it already leads Terminal-Bench 2.0 and Agents' Last Exam per the Medium reviewer — this is dismissed as "bounded" risk without considering that a single-snapshot resolution mechanism is more vulnerable to short-term leapfrogging than the forecasts' framing suggests. 5. Both forecasts note "no confirmed frontier competitor release... for September 2026" as supportive of Yes, but the brief's own timeline shows major releases occurring roughly monthly (Opus 4.8 in May, Fable 5 in June, Sonnet 5 in June, Opus 5 in July) — absence of *confirmed* announcements is weak evidence given this cadence, and a surprise competitor release (Gemini, Grok, DeepSeek) before Sept 30 is not meaningfully priced into either estimate. 6. Forecast 1's stated confidence (0.60) sits oddly below its output probability spread's implied certainty (86/14 split), and neither forecaster explains how their stated confidence level maps to the width of their probability estimate — this internal inconsistency between confidence score and forecast conviction is unaddressed by both.
6. Aggregator Opus · final calibrated forecast
82% (64% confidence)
Yes: 82% No: 18%
Anthropic holds both #1 and #2 on the exact resolution leaderboard (Agent Arena) with a ~2pp margin over Kimi K3, leads SWE-bench Verified at 96%, and has shipped four frontier models since May 2026 — all strongly favoring Yes at a Sept 30 snapshot only ~5 weeks out. The same-question Polymarket twin at 84.5% (uptrending) is the best available anchor absent a Kalshi direct price, but its thin $25.7k volume warrants a modest discount rather than trading above it. The devil's advocate correctly notes three under-weighted downside channels: (a) the resolution uses the 'Models' filter while the evidence cites the Pareto view, creating criteria-mismatch risk; (b) live regulatory/injunction dynamics already forced one Arena removal of a #1-ranked Anthropic model in June 2026 and could recur; and (c) a single-snapshot mechanism is more vulnerable to a surprise release (OpenAI already leads Terminal-Bench 2.0 and Agents' Last Exam) given the observed ~monthly frontier cadence. These push me slightly below both forecasts and marginally below Polymarket, to 82% Yes.
Pipeline Timing
Total pipeline time: 116.1s
Per-tool research timings shown in the Research section above.