# Current state
The market resolves on whichever company's model tops the LMArena text leaderboard (style-control off) on 2026-12-31. As of late July 2026, Anthropic's Claude Opus/Fable-tier models appear to hold or contest the #1 spot in several (unreliable, SEO-quality) trackers, but Google (delayed Gemini 3.5 Pro/teased "Gemini 4") and OpenAI (GPT-5.6 family) have imminent flagship releases pending, and the leaderboard has historically rotated leaders roughly monthly. No direct, verified lmarena.ai leaderboard snapshot was retrieved in this research pass.
# Timeline of key events
- 2024-09: Chatbot Arena rebrands to LMArena, moves to dedicated domain (confirmed, Wikipedia).
- 2025-04: LMArena incorporates as independent company; Llama 4 Maverick leaderboard-gaming controversy prompts policy changes (confirmed, Wikipedia).
- 2025-05: LMArena raises $100M seed at $600M valuation (confirmed, Wikipedia).
- 2026-01-06: LMArena raises $150M Series A at $1.7B valuation (confirmed, TechCrunch).
- 2026-05 (snapshot): Claude Opus 4.6 reported #1 at 1418±8, Gemini 3.1 Pro 1406, GPT-5.2 1402 — overlapping CIs, effectively tied (reported, low-reliability aggregator).
- 2026-06-09: Claude Fable 5 (Mythos-tier) launches (reported, multiple aggregators); briefly restricted/suspended (~19 days) reportedly over export-control issue (rumored, unverified).
- 2026-06-24: Gemini 3.5 Pro release reported slipping to July (confirmed via Business Insider/GDELT).
- 2026-07-01: Claude Fable 5 reportedly restored after suspension (rumored).
- 2026-07-08: Grok 4.5 launches; Musk claims parity with "last year's Claude Opus" (confirmed release, claim rumored).
- 2026-07-09: GPT-5.6 family (Luna/Terra/Sol) begins broad rollout (confirmed, GDELT/multiple outlets).
- 2026-07-16: Kimi K3 (Moonshot AI) released, largest open-weight model announced; separately accused of illegally distilling Claude (reported).
- 2026-07-19: Alibaba Qwen3.8 claims second only to Claude Fable 5 (reported).
- 2026-07-21: Google ships Gemini 3.6 Flash / 3.5 Flash-Lite / Flash Cyber, teases "Gemini 4" with no date (confirmed).
- 2026-07-24: Anthropic launches Claude Opus 5, a cheaper near-Fable-5 model (confirmed, multiple outlets).
- 2026-07-27: arena.ai leaderboard snapshot dated with 7.5M votes, 381 models, but no-style-control top-5 not retrieved (unverified).
- 2026-07-29/30: OpenAI cuts GPT-5.6 Luna price 80% amid Chinese competition (confirmed).
# Event
Will Anthropic have the best AI model (per LMArena text leaderboard, style control off) at end of December 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No direct Kalshi price was returned (kalshi_direct not populated; kalshi_related returned only an unrelated swimsuit-cover market). The primary tradable anchor available is **Polymarket, same ticker: YES = 70.5%**, up from 54% low, +7.5% over 7 days, +2% over 30 days, on modest volume ($50.7K total, 58 data points). This is a thin, single-market signal — treat with caution given no Kalshi cross-check.
# Sub-question answers
1. **Current #1 and margin** — Unclear/contested; low-reliability trackers place Claude Opus/Fable variants at or near #1 in mid-2026 snapshots (e.g., 1418 vs Gemini 1406 vs GPT-5.2 1402, essentially tied within error bars). No authoritative live lmarena.ai pull was obtained (claude_news, unverified).
2. **Has Claude ever held #1** — Multiple aggregators claim yes (Opus 4.6/4.7/4.8, Fable 5) through 2026, but source quality is poor; no Wikipedia/primary confirmation of exact historical rank achieved.
3. **Turnover frequency** — One aggregator (benchlm.ai, low reliability) reports 18 "crown changes" over 39 months — implying leadership changes roughly every ~2 months on average; broad consensus across sources is "monthly leapfrogging," no single lab sustains #1 long.
4. **2026 release timelines** — Anthropic: Opus 5 (Jul 24), Fable 5 (Jun 9, Mythos-tier). Google: Gemini 3.5 Pro delayed repeatedly (to Aug+), Gemini 4 teased without date. OpenAI: GPT-5.6 family rolled out Jul 9. xAI: Grok 4.5 (Jul 8), Musk claims parity only with "last year's" Claude Opus — implies xAI trailing frontier.
5. **Sibling markets** — polymarket_related found **zero** matching sibling markets; the code_execution tool's "OpenAI 42% / Google 27% / Anthropic 19% / xAI 7%" breakdown appears to be **synthetic/illustrative, not sourced from live data** — should be treated as unverified, not evidence.
6. **Does Anthropic deprioritize Arena-style evals** — No direct evidence found either way in this research pass; Anthropic's public benchmark emphasis (Frontier-Bench, GDPval, ARC-AGI, coding) suggests focus on agentic/coding benchmarks rather than chat-preference Arena, which could structurally disadvantage style-uncontrolled human-preference ranking, but this is inference, not confirmed fact.
# Key facts (high-confidence, factual)
1. [Wikipedia] LMArena is Berkeley-origin, now independent, VC-backed ($1.7B valuation Jan 2026).
2. [Wikipedia] Anthropic valued ~$965B (May 2026), largest pure-play AI company.
3. [GDELT/multiple] Claude Opus 5 launched Jul 24, 2026; GPT-5.6 rolled out Jul 9, 2026; Grok 4.5 Jul 8, 2026; Gemini 3.5 Pro delayed past July 2026.
4. [Wikipedia] Claude models face US federal usage restrictions (DoD "supply chain risk" designation, later enjoined) — unrelated to Arena ranking but signals enterprise/government friction.
# Cross-market signals
- Kalshi related: no matching market found.
- Polymarket (this ticker): 70.5% YES, rising trend, thin volume.
- Sibling "best AI model" categorical Polymarket group: not found/confirmed (0 matches); any percentage breakdown circulating is unverified/synthetic.
- No sportsbook signal.
# Analyst opinions and speculation
- Aggregator consensus: 2026 landscape favors "frequent leapfrogging," no lab holds #1 for a full year (felloai.com, techiehub.blog, medium.com — all low-reliability but convergent).
- Anthropic seen as currently strong/competitive but not uniquely dominant; Google and OpenAI both have unreleased flagships (Gemini 3.5 Pro/4, further GPT-5.x) that could flip leadership before December.
# Directional lean per outcome
- **Yes (Anthropic)**: Recent momentum (Opus 5, Fable 5/Mythos tier), repeated aggregator claims of #1 in mid-2026, rising Polymarket price (54%→70.5%). Opposing: historical ~2-month average tenure at #1, Google/OpenAI major releases still pending, poor source reliability undermines confidence in current Anthropic lead.
- **No (other company)**: Google's Gemini 3.5 Pro/4 and OpenAI's GPT-5.6 successors are imminent and unproven; base-rate turnover suggests high chance leadership changes hands before Dec 31; xAI/Chinese labs (Kimi K3) also closing gap, adding uncertainty against any single incumbent.
# Gaps / unknowns
- No live/authoritative lmarena.ai leaderboard pull (top-5 with scores) obtained.
- Sibling Polymarket "best model" group not actually located — percentages used elsewhere are unverified.
- No Kalshi-native price data at all for cross-check.
- Source quality throughout (SEO aggregators) is weak; many named models (Claude Fable 5, Mythos, Opus 4.6-4.8) are unverifiable against Anthropic's actual naming conventions.
# Calibration anchors
- Polymarket YES price (this ticker): 70.5%, uptrending — primary anchor given no Kalshi data.
- Historical base rate: LMArena leadership changes ~18 times/39 months (reported), implying no single-company tenure reliably spans 5+ months, let alone to year-end from current point.
- Simple model-based estimates (steady-state/Markov, base-rate) cluster around 19-33% for Anthropic — notably lower than the 70.5% market price, flagging a significant reconciliation gap between market sentiment and structural/base-rate reasoning.