← Back to scans

Will Anthropic have the best AI model at the end of September 2026?

0x537df24dbae425e973a421983b6ee570d582cac06b5fdbebef099d3d07913985 · Science and Technology · 2026-08-17
82%
Agent
88%
Market Price
-6.5%
Edge
59%
Confidence
Volume: 42,412
Spread: 1.0c
Days to resolution: 44
Markets in event: 32
Final Rationale
The best available anchor (Polymarket 88.5% on the identical ticker, trending up from ~77.5%) aligns with fundamentals: Anthropic's Claude Fable 5/Opus 5 holds or contests #1 on Arena Text Overall, while Google's Gemini 4 is only in pre-training, GPT-6 has no date, and Grok 5 has slipped twice. However, the critique's points are partly valid: the anchor is not a native Kalshi quote, the Elo cluster is <20-30 points, Opus 5's official Arena rank was unconfirmed at research time, all leaderboard data is secondhand from aggregators rather than arena.ai itself, and roughly six weeks remain in which an OpenAI point release could flip a razor-thin lead. Balancing the discount for flip/measurement risk against the market's rising, liquid signal, I land modestly below the anchor at 82% Yes, effectively agreeing with the consensus range rather than moving beyond it, since the volatility base rate cited comes from a more contested coding sub-leaderboard and no rival flagship launch is confirmed before close. Dark-horse Chinese or surprise entrants are implicitly included in the 18% No but sized near zero.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 17$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-10 75% 82% 38%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price for Anthropic in this 'best AI model end of September 2026' group, and how has it moved?
  2. What are the current prices for the sibling outcomes (Google, OpenAI, xAI, DeepSeek, Meta, Other) in the same event group, and what does the de-vigged distribution look like?
  3. Which model and company currently hold rank #1 on the LMArena Text Arena Overall leaderboard (style control off), and by what score margin?
  4. What is the highest rank Anthropic's Claude models have ever achieved on the LMArena text overall leaderboard, and how often has Anthropic held #1 historically?
  5. What new frontier model releases from Anthropic, Google, OpenAI, and xAI are expected between now and end of September 2026?
  6. How volatile is the LMArena #1 spot — how frequently has the top company changed over the past 12-18 months?
Planner reasoning
This is a Polymarket multi-outcome group (which company tops LMArena Text Arena Overall on 2026-09-30), so the direct market price plus the prices of sibling outcomes (Google, OpenAI, xAI, etc.) are the primary anchor and must sum near 1. The key empirical drivers are: where Anthropic's best model currently ranks on LMArena, Anthropic's historical track record of ever holding #1 (Claude models have typically ranked below Gemini/Grok/GPT on Arena due to Arena's human-preference bias), and expected release cadence before end-September 2026.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.2s 1 ## This Market's Polymarket Data **Will Anthropic have the best AI model at the end of September 2026?** - Current price (probability): 88.50% - 7-day price change: +8.50% - 30-day price change: +2.50% - Total volume: $42,412 (USD notional) - Price range: 77.50% - 89.50% - Data points: 28 days
polymarket_related OK 2.9s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'best AI model': 0 markets | keyword 'arena rank': 0 markets | keyword 'Anthropic': 0 markets | keyword 'Google best AI model': 0 markets | keyword 'OpenAI best AI model': 0 markets
kalshi_related OK 2.8s 2 2 related markets / summaries. keyword 'best AI model': ok | keyword 'LMArena': no matches | keyword 'Anthropic': ok
claude_news OK 32.0s 8 **Findings on LMArena / Arena.ai Text Arena Overall Leaderboard (as of mid-August 2026):** - **Current #1:** Anthropic's Claude models hold the top of the leaderboard. Multiple trackers report "Claude Fable 5 is back at #1 (~1525 ELO) after its July 1 restoration and a July 12 score re-baseline, a
claude_news OK 24.9s 16 ## Key Findings **Current state (mid-August 2026): Anthropic holds the lead across most independent benchmarks** - On BenchAlign's leaderboard, Claude Mythos 5 holds the top spot on the BenchAlign leaderboard at 83.04 overall, with Claude Fable 5 following at 82.79, and Claude Opus 5 at 82.59 . (g
gdelt_news OK 203.9s 20 GDELT: 20 articles across 3 queries (lookback=60d). 'LMArena leaderboard top model': 10 hits | 'Claude Anthropic arena leaderboard rank': 10 hits | 'Gemini tops LMArena': error GDELT rate-limited after retries (429)
wikipedia OK 0.1s 4 Fetched 4 Wikipedia entries (0 missing pages).
code_execution OK 21.3s 0 **Findings (based on manually input, illustrative assumptions since live Polymarket data wasn't pulled):** - **Assumed raw Polymarket prices** for "best AI model end of Sept 2026": Anthropic 22%, OpenAI 35%, Google 30%, xAI 8%, Meta 3%, Other 2% — these already sum to 1.00 (no overround in this ill
3. Evidence Brief Sonnet · 8201 chars
# Current state No direct Kalshi order-book snapshot was returned for this ticker; the only direct market price available is from Polymarket for the identical ticker, which prices Anthropic "Yes" at 88.5%. The underlying resolution source (arena.ai Text Arena Overall) currently shows Anthropic (Claude Fable 5 / Opus 5 cluster) holding or contesting the #1 rank as of mid-August 2026, in a tight cluster with OpenAI's GPT-5.5/5.6 family and Google's Gemini 3.1 Pro. # Timeline of key events - 2026-02-05: Claude Opus 4.6 reaches GA; Google's own Vertex docs call it "the world's best model for coding" (confirmed, per Google docs). - ~2026-03: Grokipedia snapshot shows Claude Opus 4.6 at #1 on LMArena text overall (Elo 1504), Gemini 3.1 Pro Preview #2 (1500) (reported). - 2026-06-09: Anthropic announces Claude Fable 5, becomes new frontier model (reported). - 2026-06-24: Google's Gemini 3.5 Pro release reportedly slips to July (reported). - 2026-07-01/07-12: Claude Fable 5 "restored" to #1 on Arena after a score re-baseline (~1525 Elo) (reported, third-party trackers). - 2026-07-17: Chinese Kimi K3 claims to beat Claude/GPT on a coding benchmark (reported, non-Arena benchmark). - 2026-07-19/08-03: Alibaba's Qwen3.8/Qwen3.8-Max launched, claims "second only to Claude Fable 5" (reported). - 2026-07-21: Google confirms Gemini 4 pre-training started; no model/API yet (confirmed via Google statement, per felloai). - 2026-07-24: Anthropic releases Claude Opus 5, a step-change at the Opus tier; Arena ranking not yet confirmed at release, ratings typically take 1-2 weeks to stabilize (reported). - 2026-08-01–08-16: Multiple independent trackers (localaimaster, swfte, gmicloud, felloai, modelgrep, sevenlab) show Anthropic (Claude Fable 5/Opus 5/Mythos 5) leading or near-leading most leaderboards, with one snapshot (punku.ai, Aug 7) showing GPT-5.6 Sol narrowly ahead by 0.9 points on a different composite index (reported, mixed). # Event Will Anthropic own the #1-ranked model on the arena.ai Text Arena (Overall) leaderboard at check time on 2026-09-30 12:00 PM ET? # Outcomes to forecast Yes (Anthropic #1) / No (any other company #1, including Other) # Kalshi market anchor No kalshi_direct price was retrieved for this ticker in the research pulled. The only direct price for this exact ticker comes from Polymarket: **YES 88.5%**, up +8.5% over 7 days, +2.5% over 30 days, range 77.5%–89.5% over 28 days, volume ~$42.4k. Given format overlap (identical ticker string across "Kalshi" framing and Polymarket tool), treat this 88.5% as the best available anchor for the binary Yes/No framing, but flag material uncertainty since it's not a native Kalshi quote. # Sub-question answers 1. **Polymarket price for Anthropic (this Yes/No market)** — 88.5% current, +8.5% (7d), +2.5% (30d); rising trend over past week. [polymarket_direct] 2. **Sibling outcome prices (Google/OpenAI/xAI/etc.)** — No real data found; polymarket_related found 0 matching group markets. The code_execution tool fabricated illustrative numbers (Anthropic 22%, OpenAI 35%, Google 30%, etc.) explicitly labeled as "assumed/illustrative, not live data" — **must be disregarded as non-factual**. 3. **Current #1 on LMArena Text Arena Overall** — As of mid-August 2026, Anthropic's Claude Fable 5 (~1508-1525 Elo) or Claude Opus 5 holds or is very near #1, in a tight cluster (<20-30 Elo points) with GPT-5.5 Pro/GPT-5.6 and Gemini 3.1 Pro Preview. One competing composite (punku.ai, Aug 7) shows GPT-5.6 Sol narrowly ahead by <1 point on a different index, not confirmed as the official Arena metric. [claude_news, multiple aggregators] 4. **Anthropic's historical highs on LMArena** — Held #1 with Claude Opus 4.6 (Elo 1504) as of March 2026; continued holding #1 through mid-2026 with Fable 5/Opus 5. Coding-category tracker (benchlm.ai) reports 15 "crown changes" in that sub-leaderboard over 23 months, indicating high volatility even when Anthropic often wins. [claude_news/grokipedia/benchlm] 5. **Expected frontier releases through Sept 2026** — Google's Gemini 4 is in pre-training (started July 21, 2026) with no confirmed ship date; analyst consensus points to Nov/Dec 2026 (after close date). OpenAI's GPT-6 has no confirmed date; Altman denied a Dec-2026 rumor. xAI's Grok 5 has slipped repeatedly (missed Q1 and Q2 2026 targets), unlikely by Sept 2026. Anthropic itself has the most active cadence (Fable 5 June, Opus 5 July) and may release further updates before Sept 30. [claude_news] 6. **Volatility of #1 spot (12-18mo)** — Described as "genuinely contested," with frequent crown changes especially in specialized sub-leaderboards (15 changes in coding over 23 months per benchlm.ai); overall text leaderboard shows Anthropic dominant but margins narrow (tens of Elo points) versus Google/OpenAI. [claude_news] # Key facts (high-confidence, factual) 1. [polymarket_direct] Polymarket prices Anthropic Yes at 88.5%, up sharply over the past week. 2. [Wikipedia/Claude] Anthropic released Claude Mythos and Claude Fable in 2026 as new model tiers; Opus 5 released July 24, 2026. 3. [Google Vertex AI docs, confirmed] Claude Opus 4.6 (GA Feb 5, 2026) was marketed by Google itself as best-in-class for coding/enterprise. 4. [felloai/claude_news] Google's Gemini 4 has started pre-training but has no release date before Sept 2026; most estimates place release in Q4 2026 or later. 5. [felloai/claude_news] xAI's Grok 5 has missed two prior target windows (Q1, Q2 2026) and is not expected imminently. # Cross-market signals - Kalshi related: "OpenAI or Anthropic IPO first" market prices Anthropic at 93% (context on Anthropic momentum, not model quality). "Anthropic sector classification" market at 85% (unrelated to model ranking). - Polymarket: 88.5% YES for Anthropic on this exact ticker, trending up; no sibling-outcome markets found via search. - Sportsbook implied: N/A — no sportsbook markets exist for this event. # Analyst opinions and speculation - Multiple third-party leaderboard aggregators (localaimaster, swfte, felloai, benchlm, gmicloud, sevenlab, modelgrep) converge on Anthropic leading as of mid-Aug 2026, though these are SEO/aggregator sites, not the official arena.ai source — treat as directionally suggestive, not conclusive. - One outlier snapshot (punku.ai) shows GPT-5.6 Sol nominally ahead on a different composite metric by under 1 point — underscores how thin margins are. - Consensus view: Anthropic's lead is real but fragile given monthly point-release cycles from OpenAI/Anthropic and rapid Chinese model progress (Qwen3.8, Kimi K3) narrowing gaps on adjacent benchmarks (not Arena itself). # Directional lean per outcome - **Yes (Anthropic)**: Supported by current Arena leadership (Fable 5/Opus 5), aggressive release cadence, no Google/xAI flagship expected before close, and steeply rising Polymarket price (77.5%→88.5% in a month). Opposed by: extremely tight Elo clusters (Google/OpenAI within ~20-30 points), history of frequent crown changes in adjacent sub-leaderboards, and Opus 5's Arena ranking still unconfirmed at time of research. - **No (other company)**: Supported by tight competitive clustering and possibility of OpenAI GPT-5.6/5.7 point release surprising upward; opposed by lack of any confirmed imminent Google/xAI frontier release and Anthropic's sustained multi-month Arena presence at #1. # Gaps / unknowns - No native Kalshi YES price was retrieved (tool gap) — anchor relies on Polymarket cross-price for identical ticker. - No verified sibling-outcome (Google/OpenAI/xAI) market prices; fabricated code_execution numbers must be disregarded. - Claude Opus 5's official Arena rank/score not yet confirmed at time of research (as of mid-Aug 2026). - No live arena.ai leaderboard pull was performed — all Arena rank data is secondhand via aggregator sites. # Calibration anchors - Polymarket cross-price (same ticker) as anchor: **88.5% YES**, up from ~77.5%-80% a month ago. - Historical precedent: Anthropic has held #1 on LMArena text-overall for much of 2026 (since at least March), but adjacent leaderboards (coding) show ~15 crown changes in 23 months — high underlying volatility despite recent Anthropic dominance.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.60
Yes 85%
No 15%
The only direct price for this exact ticker (Polymarket, 88.5% YES, up from ~77.5% a month prior) is the best anchor, and it aligns with the fundamentals: Anthropic's Claude Fable 5/Opus 5 holds or contests #1 on Arena Text Overall per multiple aggregators, while Google's Gemini 4 is only in pre-training (Q4 2026+ estimates) and xAI's Grok 5 has slipped twice. The main downside risk is the extremely thin Elo margins (<20-30 points versus GPT-5.5/5.6 and Gemini 3.1 Pro), one outlier composite showing GPT-5.6 Sol narrowly ahead, and the ~6-week gap between the mid-August evidence and the Sept 30 check time, during which an OpenAI point release could flip a razor-thin lead. I shade slightly below the 88.5% anchor to reflect that time-decay/flip risk and the fact that the anchor is a cross-market rather than native Kalshi quote, plus unverified official Arena data.
gpt-5.5
0.58
Yes 82%
No 18%
Sub-question evidence points toward Yes because the best available market anchor is Polymarket at 88.5% Yes and rising, while no native Kalshi quote was retrieved. Current-leader evidence also points toward Yes because multiple trackers place Anthropic's Claude Fable 5/Opus 5 at or very near #1 on Text Arena Overall, and Anthropic has held the top spot for much of 2026. Release-timing evidence points toward Yes because Google Gemini 4, OpenAI GPT-6, and xAI Grok 5 lack confirmed pre-close launches, while Anthropic has had an active June/July release cadence. I move below the 88.5% anchor to 82% because the lead is described as a tight 20-30 Elo cluster, Opus 5's official Arena rank was not yet confirmed, and the #1 spot has meaningful volatility with possible OpenAI point releases before Sept. 30.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters anchor heavily on the 88.5% Polymarket price without adequately flagging that it's explicitly not a native Kalshi quote and may reflect different liquidity/participant assumptions than a Kalshi market would—yet both converge to 82-85%, essentially just shading slightly down from the anchor rather than independently reasoning from the Arena fundamentals, producing correlated rather than diverse estimates. 2. Neither forecast weighs the "15 crown changes in coding over 23 months" volatility base rate heavily enough against a ~6-week forward window; if crown changes occur roughly every 1.5 months in a volatile sub-leaderboard, the unconditional probability of a flip in a 6-week window before Sept 30 could be non-trivial (arguably 15-25%), which is higher than the ~15-18% "No" both assign, especially since the overall leaderboard is described as "genuinely contested" with tighter margins than coding. 3. Both treat OpenAI's GPT-5.5/5.6 point-release cadence as a modest risk but understate that OpenAI has historically shipped incremental releases faster than Google/xAI—monthly point releases could flip a <30 Elo gap multiple times before the Sept 30 checkpoint, and neither forecast explicitly models multiple flip opportunities (checked only once at close) versus a single-shot risk. 4. Neither forecaster addresses the resolution-source ambiguity: the event resolves on "arena.ai Text Arena Overall" specifically, but most evidence cited comes from third-party aggregators (localaimaster, swfte, felloai, benchlm) explicitly flagged as "not the official arena.ai source"—this measurement/source risk (the official leaderboard could differ from aggregator snapshots) is not factored into either probability. 5. Both cite the punku.ai outlier (GPT-5.6 Sol ahead by <1 point) as a minor caveat, but neither considers that this data point—being from a different composite index—could indicate blind spots in how "Overall Text Arena" is being tracked, potentially signaling a closer real race than the Anthropic-favorable aggregators suggest. 6. Neither forecast substantively discusses "Other" (a dark-horse Chinese model like Kimi K3/Qwen3.8 or a surprise entrant) even though the brief notes "rapid Chinese model progress... narrowing gaps on adjacent benchmarks"—this tail scenario is implicitly bundled into "No" but never explicitly sized, likely appropriately near-zero but unaddressed nonetheless.
6. Aggregator Opus · final calibrated forecast
82% (59% confidence)
Yes: 82% No: 18%
The best available anchor (Polymarket 88.5% on the identical ticker, trending up from ~77.5%) aligns with fundamentals: Anthropic's Claude Fable 5/Opus 5 holds or contests #1 on Arena Text Overall, while Google's Gemini 4 is only in pre-training, GPT-6 has no date, and Grok 5 has slipped twice. However, the critique's points are partly valid: the anchor is not a native Kalshi quote, the Elo cluster is <20-30 points, Opus 5's official Arena rank was unconfirmed at research time, all leaderboard data is secondhand from aggregators rather than arena.ai itself, and roughly six weeks remain in which an OpenAI point release could flip a razor-thin lead. Balancing the discount for flip/measurement risk against the market's rising, liquid signal, I land modestly below the anchor at 82% Yes, effectively agreeing with the consensus range rather than moving beyond it, since the volatility base rate cited comes from a more contested coding sub-leaderboard and no rival flagship launch is confirmed before close. Dark-horse Chinese or surprise entrants are implicitly included in the 18% No but sized near zero.
Pipeline Timing
Total pipeline time: 298.7s
Per-tool research timings shown in the Research section above.