# Current state
As of the most recent verifiable data (~August 2026), the top Arena Score on the no-style-control Text leaderboard sits around 1500-1510 (Claude Opus 4.x / Gemini 3.x cluster), roughly 40-50 points short of the 1550 threshold with under 5 months remaining before the Dec 31, 2026 deadline. Progress has visibly decelerated since Nov 2025, contradicting an earlier accelerating trend.
# Timeline of key events
- 2025-03: Gemini 2.5 Pro tops leaderboard at ~1370 Elo (confirmed, lambdafin.com).
- 2025-05-16: LMArena makes "style control" the default scale, rebasing offsets to keep non-style-control comparable (confirmed, HF dataset).
- 2025-08-07: GPT-5 launches, hits "highest Arena score to date," ~1430-1442 Elo on text (confirmed, Arena/X + LessWrong).
- 2025-11-18/20: Gemini 3 Pro tops leaderboard at ~1498-1501 Elo (confirmed/reported, multiple sources).
- 2026-01: LMArena rebrands to "Arena" (arena.ai) (confirmed, Wikipedia).
- 2026-01-13: Reported vote-pipeline overhaul (sybil/identity-leak mitigation, prompt rebalancing) causing methodology-driven score fluctuations, not capability changes (reported, low-confidence SEO source, directionally consistent with official changelog).
- 2026-02 (late): Claude Opus 4.6/4.6 Thinking and Gemini 3.1 Pro Preview cluster at ~1500-1503 (reported).
- 2026-04: Sibling Polymarket "which company first hits 1550" market prices Anthropic 44% to win, "None in 2026" 36.5% (reported, implies ~63.5% chance someone hits 1550 as of that date).
- 2026-07-08 to 07-24: Wave of flagship releases — Grok 4.5, GPT-5.6 family, Kimi K3, Gemini 3.6 Flash, Claude Opus 5 (July 24) — none reported to break past ~1510 on qualifying Text Arena (reported).
- 2026-08: Top Text Arena score ~1510 (Claude Opus 4.8 per one source), with "three models above 1500" but well short of 1550; separate coding sub-leaderboard already crossed 1550-1561 (Opus 4.6) but does NOT count for this market (reported/mixed confidence).
- 2026-08-13/14: Gemini 3.7 Flash launched three weeks after prior release; Google's flagship Gemini 3.5 Pro still delayed with no timeline (confirmed, Ars Technica/GDELT).
# Event
Will any company's model reach a 1550+ Arena Score on the LMArena Text Arena (no style control) leaderboard by Dec 31, 2026?
# Outcomes to forecast
Yes (no company hits 1550 in 2026) / No (some company does hit 1550)
# Kalshi market anchor
No kalshi_direct data was returned for this ticker. The closest available anchor is the Polymarket-listed version of the identical market: "Yes" (no company hits 1550) is trading at 80.5%, up +2.5% over 7 days, +0.5% over 30 days, range 52%-81.5% over 90 days, but thin volume (~$25k total). Treat as a proxy anchor, not a confirmed Kalshi print.
# Sub-question answers
1. **Current top score/model** — As of ~Aug 2026, ~1500-1510, led by Claude Opus 4.x variants (Opus 4.6/4.8) with Gemini 3.x close behind; low-confidence SEO sources, but directionally corroborated by multiple outlets (claude_news).
2. **Score growth per quarter** — ~1370 (Mar'25) → ~1435 (Aug'25) → ~1500 (Nov'25) → ~1510 (Aug'26): roughly +65pts/8mo through late 2025, then only +10pts over the following ~9 months — a sharp deceleration, not the acceleration a naive linear fit suggests (code_execution vs. claude_news conflict; the news trajectory is more granular/recent and should be weighted higher).
3. **2026 frontier releases** — GPT-5.6 (Jul), Grok 4.5 (Jul), Gemini 3.6/3.7 Flash (Jul/Aug), Claude Opus 5 (Jul 24); Gemini 3.5 Pro flagship delayed indefinitely; GPT-6, Grok 5, Gemini 4, Claude Opus 6 all unconfirmed/rumored only as of Aug 2026 (claude_news).
4. **Scale recalibration** — Yes: style-control became default (May 2025) with rebasing; a Jan 2026 vote-pipeline overhaul reportedly caused fluctuations unrelated to capability — both inject noise into cross-period score comparisons (HF dataset, reported source).
5. **Sibling markets** — April 2026 per-company Polymarket legs implied ~63.5% chance someone hits 1550 (Anthropic 44%, None 36.5%); by inference, current pricing has shifted toward "None," consistent with the 80.5% "Yes" (no-hit) price now observed.
6. **Related Kalshi/Polymarket cross-signal** — No other AI-benchmark markets found on Kalshi (kalshi_related turned up only an unrelated Swimsuit Issue market); Polymarket_related found zero matching markets, limiting cross-venue triangulation.
# Key facts (high-confidence, factual)
1. [claude_news/lesswrong] GPT-5 scored ~1430-1442 Elo (Aug 2025).
2. [claude_news/multiple] Gemini 3 Pro scored ~1498-1501 Elo (Nov 2025).
3. [HF dataset] Style-control became default scale May 16, 2025, with rebasing to preserve non-SC comparability.
4. [Wikipedia] LMArena rebranded to "Arena" (arena.ai) in Jan 2026.
5. [claude_news] Coding sub-leaderboard (not the qualifying category) already exceeded 1550-1561 via Claude Opus 4.6.
# Cross-market signals
- Polymarket (same market): 80.5% "Yes" (no 1550 hit), rising slightly, low volume.
- Sibling Polymarket (per-company legs, April 2026 snapshot): implied ~63.5% someone hits 1550 — stale relative to current pricing, suggesting the market has moved toward "None" since then.
- No Kalshi cross-market data found.
# Analyst opinions and speculation
- code_execution's naive linear extrapolation (using older 4-point series) implies near-certain breach of 1550 (P(No)≈0-1%), but this conflicts with more granular 2026 news data showing a plateau around 1500-1510 since Nov 2025.
- claude_news synthesis frames the pattern as "incremental leapfrogging" with top-8 models clustered within ~55 Elo points — a maturity/plateau narrative more consistent with the current 80.5% "Yes" price.
- Many 2026-dated numeric claims (SEO sites) are low-confidence and internally inconsistent (implausible model names like "Claude Opus 4.8," "Fable 5").
# Directional lean per outcome
- **Yes (no 1550 hit)**: Supported by observed deceleration since Nov 2025 (+10pts/9mo), Gemini Pro flagship delays, clustering/plateau pattern, and market pricing at 80.5%.
- **No (someone hits 1550)**: Supported by continued rapid model cadence (multiple flagship launches every 4-8 weeks), coding sub-leaderboard already surpassing 1550 (showing capability headroom exists), and ~4-5 months remaining with unconfirmed GPT-6/Grok5/Gemini4/Opus6 possibly launching.
# Gaps / unknowns
- No confirmed Kalshi YES price for this exact ticker was retrieved.
- Post-August 2026 data is entirely absent; unclear if any Sept-Dec 2026 flagship closed the ~40pt gap.
- Reliability of specific 2026 model names/scores from SEO sources is low; official arena.ai leaderboard not directly queried.
# Calibration anchors
- Polymarket proxy YES (no 1550 hit) price: 80.5% (anchor).
- Historical precedent: score jumps of 60-70 points occurred within ~3-8 month windows in 2025, but growth has since slowed to ~1pt/month pace through mid-2026 — suggests base rate for a further 40pt jump in ~5 months is uncertain but trending unfavorable for "No."