# Current state
The market resolves YES if any model on the LMArena (formerly Chatbot Arena) text leaderboard (style-control unchecked) hits ≥1550 Elo by Dec 31, 2026. As of the latest research (~Aug 2026), the top models cluster in the ~1500–1525 range (Claude Opus 4.6/4.7/4.8, GPT-5.5 Pro, Gemini 3.1 Pro), well below 1550, with third-party trackers disagreeing on the exact top figure (1500 vs. 1510 vs. 1525) and no direct lmarena.ai leaderboard snapshot retrieved.
# Timeline of key events
- 2024-12: o1-2024-12-17 reaches ~1400 Elo (top score), per benchlm.ai leaderboard history — reported.
- 2025-03: chatgpt-4o-latest-20250326 reaches ~1450 — reported.
- 2025-06: gemini-2.5-pro reaches ~1500 (first crossing of 1500 barrier) — reported.
- 2025-11: Robinhood prediction market prices "≥1500 Elo" at 12%, "≥1550" at 3%, "≥1600" at 2% for resolution before Jan 1, 2026 — confirmed market pricing, implying low probability assigned at that time.
- 2026-01-28: LMArena officially rebrands to "Arena"; scoring methodology (Bradley-Terry/Elo) unchanged — reported.
- 2026-02: claude-opus-4-6 cited as reaching ~1500 on benchlm.ai history — reported.
- 2026-07-01: "Claude Fable 5" reportedly restored to leaderboard (model name unverified/suspect) — rumored, low-confidence source.
- 2026-07-12: A leaderboard "score re-baseline" event reported by localaimaster.com — rumored, could distort trend continuity.
- 2026-08 (current): Top tier reported clustered at 1500–1525 (Claude Opus 4.6/4.7/4.8, GPT-5.5 Pro, Gemini 3.1 Pro Preview); top-10 models within ~20 Elo points of each other (toolcenter.ai) — reported, moderate confidence.
# Event
Will any AI model reach a Chatbot Arena (LMArena text leaderboard, no style control) score of at least 1550 by December 31, 2026?
# Outcomes to forecast
- Yes
- No
# Kalshi market anchor
No kalshi_direct data was returned for this ticker (only an unrelated Kalshi "AI benchmark" market on SCOTUS/ERISA, not relevant). The best available direct market pricing comes from Polymarket on the identical question: **current YES price ≈ 9.5%**, down from a 90-day high of 34.5%, down 8.5pts over 30 days and 1pt over 7 days; total volume ~$50k (thin). Treat this as the primary cross-market anchor in absence of native Kalshi pricing.
# Sub-question answers
1. **Current top score/model** — Disputed across trackers (Aug 2026): ~1500 (metatext.io, Claude Opus 4.6 Fast), ~1510+ (swfte.com, Claude Opus 4.8), ~1525 (localaimaster.com, "Claude Fable 5" — unverified model name). Consensus range: roughly 1500–1525, no confirmed source above 1525.
2. **Pace of increase** — Milestone history: ~1400 (Dec 2024) → ~1450 (Mar 2025) → ~1500 (Jun 2025) → ~1500-1525 (early-mid 2026). Quarterly gains decelerated from ~45 pts/quarter (early 2025) to ~20-30 pts/quarter (late 2025/2026) (code_execution analysis). Linear extrapolation implies 1550 crossed easily by 2026; a decelerating/saturating model implies crossing only ~mid-2027, after the deadline.
3. **Frontier releases expected** — GPT-5.5 Pro, Gemini 3.1 Pro Preview, Claude Opus 4.6/4.7/4.8 already reflected in current ~1500-1525 scores; no confirmed GPT-6 or Gemini 4 in research. Historically each frontier release has added ~20-50 Elo points, suggesting 1-2 more major releases could plausibly reach 1550, but recent releases show smaller gains (compression).
4. **Methodology changes** — Platform rebranded LMArena→Arena (Jan 28, 2026); a "score re-baseline" was reported July 12, 2026 (unverified quality source), which could artificially shift scores without capability changes — a wildcard for resolution consistency.
5. **Cross-market signals** — Polymarket YES ~9.5%, declining trend. No Polymarket-related markets found (0 matches for lmarena/arena score keywords). Kalshi has no matching native market; a Robinhood contract (Nov 2025, for Jan 1 2026 deadline) priced ≥1550 at only 3%, ≥1500 at 12% — much stricter deadline than this market's Dec 2026 date, so not directly comparable but shows historical market skepticism about rapid Elo growth.
6. **Saturation signs** — Yes: top-10 models within ~20 Elo points of each other (toolcenter.ai, May 2026); multiple trackers show tight clustering at 1500-1525, consistent with slowing marginal gains at the frontier.
# Key facts (high-confidence, factual)
1. [wikipedia] Elo/Bradley-Terry ratings are relative to the competitor pool, not absolute — score inflation/deflation possible from re-baselining or new entrants.
2. [benchlm.ai via claude_news] Top score trajectory: 1400 (Dec 2024) → 1450 (Mar 2025) → 1500 (Jun 2025).
3. [claude_news] Platform rebranded LMArena→Arena on 2026-01-28; scoring methodology unchanged.
4. [polymarket_direct] Current YES price 9.5%, down from 34.5% high over past 90 days.
# Cross-market signals
- Kalshi related: no matching native market found; unrelated ERISA market only.
- Polymarket: 9.5% YES, declining 30-day trend (-8.5pts), thin volume (~$50k total).
- Sportsbook implied: N/A (not applicable to this event type). Robinhood historical AI-capability contract (different deadline, Jan 2026) priced ≥1550 at only 3%.
# Analyst opinions and speculation
- code_execution quantitative models split widely: naive linear extrapolation → ~100% by Dec 2026; saturating/log model → ~32%; blended/Monte Carlo estimates → 65-71%. Wide range reflects uncertainty over whether 2025-26 deceleration is temporary or structural.
- Multiple low-quality aggregator blogs (localaimaster, swfte, messengerbot) reference unverified model names, reducing confidence in claims of scores already near/above 1525.
# Directional lean per outcome
- **Yes**: Historical Elo growth is rapid in raw terms (350+ pts in ~2 years); several more frontier releases (GPT-5.6+, Gemini 3.2+, Claude Opus 5) plausible before Dec 2026; re-baselining events could push scores up mechanically.
- **No**: Market pricing (Polymarket 9.5%, declining) strongly favors No; observed 2025-2026 deceleration and top-10 clustering within ~20 points suggest saturation; current top score (~1500-1525) needs +25-50 points with no confirmed model near threshold; historical analogous market (Robinhood, tighter deadline) priced similar threshold at only 3%.
# Gaps / unknowns
- No native Kalshi price for this exact ticker was retrieved — anchor relies on Polymarket for the same question.
- No direct lmarena.ai leaderboard scrape; current top score is disputed (1500-1525) across low-reliability secondary sources.
- Unclear whether "score re-baseline" (Jul 2026) is real/material or would affect resolution eligibility.
- No confirmed roadmap for GPT-6, Gemini 4, or Claude 5 within 2026.
# Calibration anchors
- Polymarket current YES price (this exact question): 9.5%, trending down.
- Robinhood analogous market (Nov 2025, shorter deadline): priced ≥1550 Elo at only 3%.
- Quantitative trend models (code_execution): wide range 32%-100%, blended estimate ~65-70%, but this conflicts sharply with live market pricing (~9.5%), suggesting market participants weight saturation/plateau evidence heavily over naive extrapolation.