← Back to scans

Will Anthropic have the best Code Arena | WebDev AI at the end of September 2026?

0xe16d6ef2442a8877d457ad2292e6f9832f569911685d5f2e7b0c89fbcf75cb7b · Science and Technology · 2026-08-27
85%
Agent
84%
Market Price
+1.0%
Edge
74%
Confidence
Volume: 25,346
Spread: 4.0c
Days to resolution: 34
Markets in event: 30
Final Rationale
Anthropic currently holds a clear #1 (Claude Opus 5, ~1691–1704 Elo, ~29 above Kimi K3) and has led this leaderboard ~19 of the past 20 months, with only one company-level hand-off (Gemini 2.5 Pro, May 2025). Rival flagships that could plausibly displace it — Gemini 3.5 Pro (still enterprise preview), GPT-5.7/Astra (unreleased/unnamed), Grok 5 (repeatedly delayed) — have no confirmed pre-Sept-30 launch, and self-dethronement by another Claude still resolves YES. The devil's advocate is right that all evidence is Jul–Aug snapshots with a ~2-month blind window, that Kimi K3's confidence interval partly overlaps Opus 5's, and that the market's 23.5%–87% range plus a +9.5pt weekly move suggests momentum rather than settled consensus — so I decline the small upward nudge both forecasters made and sit essentially at the 84% anchor. Netting the strong base rate against genuine surprise-launch and data-lag risk gives 85% YES.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 7$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-20 77% 74% 64%
2026-08-13 66% 68% 56%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related claude_news claude_news kalshi_related gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price for Anthropic (and its rivals Google/OpenAI/xAI) in this leaderboard event group?
  2. Which company's model currently occupies rank #1 on the arena.ai Code Arena | WebDev leaderboard, and by what margin?
  3. How frequently has the #1 spot on the WebDev/coding arena leaderboard changed hands over the past 12-18 months, and what is the average tenure of a leader?
  4. What frontier coding-capable model releases (Claude, Gemini, GPT, Grok) are expected or rumored between now and September 30, 2026?
  5. Has Anthropic historically topped the WebDev arena (e.g., Claude 3.7/4/4.5 Sonnet, Opus), and how does its style-controlled ranking compare to Google's Gemini and OpenAI's GPT models?
  6. Are there any structural factors (Anthropic's participation/opt-in to LMArena, arena.ai rebrand/methodology changes) that affect whether Anthropic models appear at the top?
Planner reasoning
This is a Polymarket question about which company tops the arena.ai Code Arena | WebDev leaderboard on Sept 30, 2026, so the Polymarket price across the sibling markets (Google, OpenAI, Anthropic, xAI, etc.) is the primary anchor. The key empirical drivers are: who currently leads WebDev arena, how persistent leadership has been historically, and what frontier model releases are expected before end of September 2026.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will Anthropic have the best Code Arena | WebDev AI at the end of September 2026?** - Current price (probability): 84.00% - 7-day price change: +9.50% - 30-day price change: +3.50% - Total volume: $25,346 (USD notional) - Price range: 23.50% - 87.00% - Data points
polymarket_related OK 3.3s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'Code Arena WebDev': 0 markets | keyword 'best AI model end of September': 0 markets | keyword 'LMArena leaderboard': 0 markets | keyword 'Anthropic best model': 0 markets
claude_news OK 39.3s 14 Based on available data (arena.ai's Code Arena / WebDev leaderboard, formerly LMArena's WebDev Arena): **Current standing (most recent snapshots, July–Aug 2026):** - Latest confirmed snapshot (July 31, 2026): Claude Opus 5 leads this snapshot at 1704 (±16), and it clears the confidence interval of
claude_news OK 31.3s 13 Based on my research, here are the key findings: - **Current leader (as of Aug 2026):** Claude Opus 5 tops the WebDev/Code Arena leaderboard at ~1691-1704 Elo, having dethroned the previous Anthropic model (Claude Fable 5) and beaten new entrants Kimi K3 (open-weight, ~1674-1675) and Grok 4.6 (~163
kalshi_related OK 3.2s 1 1 related markets / summaries. keyword 'best AI model': no matches | keyword 'LMArena': no matches | keyword 'Anthropic': ok
gdelt_news OK 121.6s 0 GDELT: 0 articles across 3 queries (lookback=60d). 'LMArena WebDev leaderboard Claude': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Read timed out. (read timeout=30) | 'arena.ai leaderboard top model': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retrie
wikipedia OK 0.1s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 58.4s 0 **Key quantitative findings** - **Raw market vig:** Illustrative Polymarket-style prices for the WebDev-arena "best company" group (Anthropic 47¢, Google 24¢, OpenAI 16¢, xAI 7¢, Other 9¢) sum to 1.030 → an overround of **~3.0%**. - **De-vigged probabilities:** Both normalization methods agree clo
3. Evidence Brief Sonnet · 6689 chars
# Current state Anthropic's Claude models (Opus 4.5 → 4.6 → 4.7 → Opus 5) have held #1 on the arena.ai Code Arena/WebDev leaderboard almost continuously since Nov 2025, and per multiple Aug 2026 snapshots Claude Opus 5 currently leads at ~1691–1704 Elo, ahead of Kimi K3 (~1675) and Grok 4.6 (~1630). The market's own price (this exact ticker) sits at 84% YES, up sharply (+9.5% in 7 days) on high conviction. # Timeline of key events - 2024-12: WebDev Arena launches; Claude 3.5 Sonnet (Oct 2024) leads. (confirmed, aiwiki.ai) - 2025-02: Claude 3.7 Sonnet takes #1, Anthropic retains lead. (confirmed, aiwiki.ai) - 2025-05: Gemini 2.5 Pro (I/O edition) briefly takes #1 — only known Anthropic interruption. (reported, dev.to) - 2025-07: Grok-4 ranks only #12, well behind Claude/Gemini/GPT. (confirmed, arena.ai/X) - 2025-11: Code Arena relaunch; Claude Opus 4.5 takes #1, "surpassing Gemini 3 Pro"; top-5 dominated by Anthropic (Opus 4.1, Sonnet 4.5 x2), then GPT-5, GLM-4.6. (confirmed, arena.ai/X) - 2026-02: Claude Opus 4.6 pulls visibly clear of a crowded field of 38+ models. (reported, kearai.com) - 2026-05-24: Claude Opus 4.7 variants occupy top slots (5 of top 10 = Claude); first OpenAI entry outside top 10. (reported, aiwiki.ai/propelcode) - 2026-07-08 / 2026-08-12: xAI ships interim Grok 4.5 then Grok 4.6 (not Grok 5); Grok 4.6 lands only #3 (~1630 Elo). (confirmed, geotoolbox.ai) - 2026-07-31: Claude Opus 5 becomes new WebDev Arena leader (~1691–1704 Elo), dethroning Anthropic's own Claude Fable 5; Kimi K3 (#2, open-weight) trails by ~29 Elo. (confirmed, ainexhub.com) - 2026-08-18: GPT-5.7/"Astra" (OpenAI's next frontier model) still unreleased, name undecided; Gemini 3.5 Pro missed rumored mid-July release, remains enterprise-preview only. (reported, orcarouter.ai / cometapi.com) # Event Will Anthropic own the #1-ranked model on arena.ai's Code Arena | WebDev leaderboard when checked Sept 30, 2026, 12:00 PM ET? # Outcomes to forecast Yes (Anthropic #1) / No (any other company #1) # Kalshi market anchor Current YES price (this ticker, via Polymarket-style data feed): **84%**, up from ~74.5% a week ago (+9.5% 7d) and ~80.5% 30d ago; volume $25,346 total; historical range 23.5%–87%, suggesting the market recently re-priced upward toward near-certainty after some earlier volatility (low of 23.5% likely reflects an early/uninformed period). # Sub-question answers 1. **Polymarket pricing for rivals** — No separate per-company Polymarket markets found (polymarket_related returned 0 matches); only the Anthropic YES/No contract itself is tracked, at 84%. A code_execution tool fabricated illustrative multi-company prices (Anthropic 47¢, Google 24¢, OpenAI 16¢, xAI 7¢) but these are **synthetic/illustrative, not real market data** and conflict with the actual 84% Anthropic price — disregard as speculative. 2. **Current #1** — Claude Opus 5 (Anthropic) leads at ~1691–1704 Elo as of the July 31–Aug 2026 snapshots, clearing the #2 confidence interval (Kimi K3, ~1675–1674, open-weight) by ~29 Elo — described as "a rare clear #1." (claude_news/ainexhub.com) 3. **Turnover frequency** — Over ~20 months (Dec 2024–Aug 2026), the #1 spot changed company hands only once (Google's brief May 2025 Gemini 2.5 Pro stint); otherwise Anthropic has held #1 continuously, implying long average tenure (many months) and low base-rate hazard of near-term hand-off. 4. **Upcoming releases** — Gemini 3.5 Pro rumored mid-July 2026, missed that window, remains in limited enterprise preview only (as of Aug 2026). GPT-5.7/"Astra"/GPT-6 unreleased and unnamed as of Aug 18, 2026. Grok 5 repeatedly delayed (targeted late 2025→Q1→Q2 2026, all missed); xAI shipped interim Grok 4.5/4.6 instead, with Grok 4.6 landing only #3. 5. **Historical Anthropic performance** — Claude has led WebDev Arena almost continuously since Dec 2024 (3.5 Sonnet → 3.7 Sonnet → Opus 4.1/4.5/4.6/4.7 → Opus 5), with 5 of top 10 slots Claude variants in a May 2026 snapshot. Google/OpenAI have intermittently reached #2 (Gemini 2.5 Pro won briefly; GPT-5/5.2 reached #2 at times) but not sustained #1. 6. **Structural factors** — No evidence Anthropic has opted out of arena.ai/Code Arena; it actively participates and ships new Opus variants there. LMArena→arena.ai rebrand (Nov 2025 "Code Arena" relaunch) did not disrupt Anthropic's lead — it took #1 immediately post-relaunch. # Key facts (high-confidence, factual) 1. [polymarket_direct] Current YES = 84%, +9.5% 7d, volume $25,346. 2. [claude_news/arena.ai] Claude Opus 5 currently #1 (~1691–1704 Elo), clear margin over Kimi K3. 3. [aiwiki.ai] Anthropic led continuously Dec 2024–present except brief May 2025 Gemini interruption. 4. [multiple] Grok 5 and GPT-5.7/Astra unreleased as of Aug 2026; Gemini 3.5 Pro delayed to limited preview. # Cross-market signals - Kalshi related: "OpenAI or Anthropic IPO first — Anthropic" at 92% (unrelated but shows market confidence in Anthropic generally); no direct Kalshi arena market found. - Polymarket: only this ticker exists for the question; no rival-company sub-markets to cross-check. - Sportsbook implied: N/A. # Analyst opinions and speculation - llm-stats.com/kearai.com commentary treats Claude Opus 5's lead as the "most discriminating signal at the frontier," implying analyst consensus favors continued Anthropic dominance. - Some sources note the field below #1 is "the most competitive landscape in AI coding history," suggesting fragility of any single-leader gap even as Anthropic currently sits clear. # Directional lean per outcome - **Yes (Anthropic)**: Strong — 20-month near-uninterrupted leadership, current clear Elo lead, rivals' flagship releases (Gemini 3.5 Pro, GPT-5.7/6, Grok 5) all delayed with no confirmed pre-Sept-30 launch, market price already high (84%) and trending up. - **No**: Weak but nonzero — only ~5 weeks to close; a rival could still ship (Gemini 3.5 Pro enterprise preview could go GA; Grok 5 "surprise" launch); one historical precedent (Gemini 2.5 Pro, May 2025) shows hand-offs are possible even amid Anthropic strength; open-weight Kimi K3 closing gap (~29 Elo) faster than prior challengers. # Gaps / unknowns - No live confirmation of leaderboard state exactly at Sept 30, 2026 check time (data is Jul–Aug 2026 snapshots). - No genuine per-company Polymarket sub-markets to triangulate against the 84% price. - Fabricated/illustrative multi-company pricing from code_execution should not be used as evidence. # Calibration anchors - Kalshi/Polymarket current YES price (anchor): **84%** - Historical base rate: Anthropic held #1 ~19 of past ~20 months (~95% uptime) on this leaderboard; only 1 confirmed hand-off in that period.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.72
Yes 87%
No 13%
Claude Opus 5 currently holds a rare clear #1 on Code Arena/WebDev at ~1691-1704 Elo, ~29 Elo ahead of Kimi K3, and Anthropic has held the top spot in ~19 of the past 20 months with only one company-level hand-off (Gemini 2.5 Pro, May 2025). The main challengers' flagships — Gemini 3.5 Pro (still enterprise-preview after missing mid-July), GPT-5.7/Astra (unreleased and unnamed as of Aug 18), and Grok 5 (repeatedly delayed; interim Grok 4.6 only #3) — have no confirmed launch before Sept 30, and even a launch would need to clear a sizable Elo gap. The Kalshi/Polymarket anchor at 84% and trending up (+9.5% in 7 days) is consistent with this; I sit slightly above it given the short ~5-week window, the low historical hazard rate, and Anthropic's own habit of dethroning itself (which still resolves YES). Residual risk comes from data being Jul-Aug snapshots rather than live, plus surprise-release tail risk.
gpt-5.6-sol
0.76
Yes 85%
No 15%
The 84% YES market price is the primary anchor, and the current leaderboard points toward Yes because Claude Opus 5 has a roughly 29-Elo lead over Kimi K3. Turnover history and Anthropic's sustained performance also favor Yes: Anthropic held first place for about 19 of the past 20 months and currently has several highly ranked variants. Upcoming-release evidence favors Yes because Gemini 3.5 Pro, GPT-5.7/Astra, and Grok 5 remain delayed or unreleased, while the latest Grok reached only third place. I move slightly above the market to 85%, while retaining meaningful No probability because Kimi is close and a surprise rival launch could alter the ranking before September 30.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters treat Anthropic self-dethronement (e.g., Opus 5 succeeded by another Anthropic model) as automatically resolving YES, but neither flags that open-weight Kimi K3 closing to ~29 Elo is unusually fast relative to prior challengers, per the brief's note that "the field below #1 is the most competitive landscape in AI coding history" — this fragility signal is underweighted in both rationales' No-side reasoning. 2. Neither forecaster explicitly addresses the data-lag gap: all Elo snapshots are from Jul 31–Aug 2026, nearly two months before the Sept 30 check, yet both treat current standings as near-decisive without much discount for unknown intervening leaderboard shuffles (new open-weight Chinese models, minor Gemini/GPT point releases) not captured in the brief. 3. Both anchor tightly around the 84% market price and nudge only slightly upward (85-87%), but the market's own volatility (23.5%–87% historical range, +9.5% in just 7 days) suggests real uncertainty/momentum-chasing rather than settled consensus — neither forecaster interrogates why the price moved 9.5 points in a week or whether that move is durable versus a temporary spike that could mean-revert. 4. The brief notes the single historical hand-off (Gemini 2.5 Pro, May 2025) occurred despite Anthropic being in a strong position at the time too — neither forecast quantifies how a "surprise" GA launch of Gemini 3.5 Pro or an unexpected Grok 5 ship (both plausible within a 5-week window per repeated pattern of surprise releases) could specifically replicate that precedent; the ~13-15% No-probability seems to reflect base rate more than an actual scenario-weighted estimate of a rival launch clearing the Elo gap. 5. Both cite "Anthropic's own habit of dethroning itself" as reducing risk, but neither considers whether Arena methodology changes (e.g., another Code Arena "relaunch" akin to Nov 2025) could introduce discontinuities that temporarily displace Anthropic during a transition, which the brief's "structural factors" section flags as a past disruption vector even though it resolved in Anthropic's favor last time. 6. Neither forecast discusses the confidence-interval overlap risk explicitly — Kimi K3's ~1675-1691 CI is described as approaching Opus 5's ~1691-1704, meaning under certain Elo-ranking methodologies the "clear" 29-point lead could be within statistical noise, a nuance that would justify a wider probability spread than the tight 85-87% both converged on.
6. Aggregator Opus · final calibrated forecast
85% (74% confidence)
Yes: 85% No: 15%
Anthropic currently holds a clear #1 (Claude Opus 5, ~1691–1704 Elo, ~29 above Kimi K3) and has led this leaderboard ~19 of the past 20 months, with only one company-level hand-off (Gemini 2.5 Pro, May 2025). Rival flagships that could plausibly displace it — Gemini 3.5 Pro (still enterprise preview), GPT-5.7/Astra (unreleased/unnamed), Grok 5 (repeatedly delayed) — have no confirmed pre-Sept-30 launch, and self-dethronement by another Claude still resolves YES. The devil's advocate is right that all evidence is Jul–Aug snapshots with a ~2-month blind window, that Kimi K3's confidence interval partly overlaps Opus 5's, and that the market's 23.5%–87% range plus a +9.5pt weekly move suggests momentum rather than settled consensus — so I decline the small upward nudge both forecasters made and sit essentially at the 84% anchor. Netting the strong base rate against genuine surprise-launch and data-lag risk gives 85% YES.
Pipeline Timing
Total pipeline time: 214.2s
Per-tool research timings shown in the Research section above.