# Current state
As of the most recent data (~August 2026), the top score on the Arena.ai (formerly LMArena) Text leaderboard, style-control off, sits in an estimated ~1510–1525 range (unverified aggregator trackers), still 25–40 points below the 1550 threshold. The 1500 mark was only reached/exceeded during 2026 (Gemini 3 Pro debuted at ~1487-1498 in Nov 2025, first to approach 1500). Polymarket prices "Yes" at 10.5%, down sharply from a 34.5% high over the past 30 days.
# Timeline of key events
- 2024-mid: GPT-4o leads at ~1290 Elo (code_execution trend fit). [confirmed/trend]
- 2025-early: Gemini 2.5 Pro leads at ~1440 Elo. [confirmed/trend]
- 2025-11-18/20: Gemini 3 Pro launches, tops leaderboard at ~1487–1498 Elo, edging Grok 4.1. [reported, claude_news]
- 2025-12-31: A sibling Manifold market on "Gemini 3 reaches 1500+ by Dec 31 2025" resolves NO, confirming score stayed <1500 through year-end 2025. [confirmed]
- 2026-01-28: LMArena rebrands to "Arena" (now at arena.ai instead of lmarena.ai). [reported]
- 2026-02-20: Anthropic Opus 4.6 (~1504) and Gemini 3.1 Pro (~1500) essentially tied at top. [reported]
- 2026-04: Claude Opus 4.6 Thinking leads at a "record" 1504 Elo. [reported]
- 2026-05: Top score reaches ~1501, held by a thinking-enabled Claude variant; top-5 spread within 20–30 points. [reported]
- 2026-07-01/12: Arena restored after outage; re-baselined scoring to count only post-restoration votes, breaking strict continuity with prior scores. [reported]
- 2026-07-26: Kimi K3 (open-weight) ships, leads a specialized Frontend Code sub-arena (not overall text). [reported]
- 2026-08: Aggregator trackers describe a tight top cluster (~1500–1525): Claude Opus 4.8/4.7, GPT-5.5/5.5 Pro, Gemini 3.1 Pro; one low-reliability source claims "Claude Fable 5" at ~1525 overall. Coding-specific sub-leaderboard already shows Opus 4.8 at ~1582 (not the resolution metric). [reported, low confidence]
# Event
Will any model reach an Arena Score ≥1550 on the LMArena/Arena.ai Text Arena overall leaderboard (style control off) by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct pricing was returned (kalshi_related found 0 matches); the Polymarket price on the identical market is the best available anchor: **YES = 10.5%**, down from a 90-day high of 34.5%, 7-day change -1.0%, 30-day change -8.5%, total volume ~$49K. Trend is clearly deteriorating for "Yes."
# Sub-question answers
1. **Current highest score/model** — As of ~Aug 2026, estimates cluster at ~1500–1525 (overall text, style-off), with Claude Opus 4.7/4.8, GPT-5.5 Pro, and Gemini 3.1 Pro tightly bunched near the top; one unverified tracker claims "Claude Fable 5" at ~1525. No single authoritative reading confirms an exact current leader/score. [claude_news, low-med confidence]
2. **12-24 month growth trajectory** — Roughly ~1290 (mid-2024, GPT-4o) → ~1440 (early 2025, Gemini 2.5 Pro) → ~1487-1498 (Nov 2025, Gemini 3 Pro) → ~1500-1504 (Feb-May 2026) → ~1510-1525 (Aug 2026). Linear fit ≈11.4 pts/month over the full window; recent segment ≈12 pts/month. [code_execution, claude_news]
3. **Sibling Polymarket thresholds (1500/1525/1550/1575/1600)** — Not returned by polymarket_related (0 matches found); only the 1550 market itself was retrieved (10.5% Yes). No cross-threshold distribution could be constructed.
4. **Expected 2026 frontier releases** — Reported/rumored iterations of GPT-5.4/5.5/5.6, Claude Opus 4.7/4.8 (and possibly Opus 5), Gemini 3.1 Pro (and potential 3.x updates), plus open-weight pushes (DeepSeek, Qwen 3.7 Max, Kimi K3) are already reflected in the Aug 2026 cluster near 1510-1525. [claude_news]
5. **Methodology changes** — Platform rebranded LMArena→"Arena" (Jan 2026, moved to arena.ai). A July 2026 outage/restoration triggered a re-baselining that counts only votes since July 1, 2026 restoration — breaking strict score continuity and adding uncertainty to comparisons with pre-2026 scores. Core Bradley-Terry/Elo voting methodology otherwise unchanged. [claude_news]
6. **Step jumps from flagship launches** — Gemini 3 Pro's Nov 2025 launch produced the biggest jump (~1440→~1490), a ~50-point move. Since then, jumps have been smaller/incremental (~5-15 pts per release; overall score moved only ~1500→~1520 over ~9 months despite multiple flagship launches — Opus 4.6/4.7/4.8, GPT-5.4/5.5, Gemini 3.1 Pro). No single 2026 release has produced a >20-point jump on the overall leaderboard. [claude_news]
# Key facts (high-confidence, factual)
1. [polymarket_direct] Current Yes price is 10.5%, down from 34.5% high in the past 90 days; 30-day trend -8.5%.
2. [claude_news/Manifold] A sibling 2025 market on Gemini 3 reaching 1500+ resolved NO, confirming <1500 through end-2025.
3. [claude_news] Coding-specific sub-arena scores have already exceeded 1550 (~1567-1582), but this is NOT the resolution metric (overall Text Arena required).
4. [code_execution] Trend-based projections to Dec 31, 2026 range widely: linear/recent-rate models imply ~1640-1656 (near-certain Yes), decelerating (sqrt) models imply ~1548 (~coin-flip), log-decel models imply low scores (~unlikely Yes).
# Cross-market signals
- Kalshi related: none found (0 matches).
- Polymarket: 10.5% Yes on this exact market, sharply declining trend (was 34.5% at some point in last 90 days) — suggests market participants have grown more pessimistic as mid-2026 progress stalled near 1500-1525 rather than accelerating toward 1550.
- Sportsbook implied: N/A (not applicable to this category).
# Analyst opinions and speculation
- Aggregator/SEO trackers (low reliability) suggest overall scores may already be nearing 1525 as of Aug 2026, which if accurate would narrow the gap to ~25 points with 4 months remaining — plausible but unconfirmed against the official arena.ai site.
- Bottom-line consensus across research tools: reaching 1550 by Dec 2026 is "plausible but not a done deal," requiring sustained ~25-40 point gains in the remaining months, roughly consistent with the pace of the past 9 months (30-40 pts) but faster than the most recent few months' incremental (~5-15 pt) releases.
# Directional lean per outcome
- **Yes**: Historical full-2-year pace (~11-12 pts/month) extrapolated linearly would clear 1550 comfortably; frontier lab release cadence remains high (multiple GPT-5.x, Claude Opus, Gemini updates in 2026).
- **No**: Growth has visibly decelerated since Gemini 3's initial jump — score moved only ~1500→~1520 despite 9 months and 5+ flagship releases (Opus 4.6-4.8, GPT-5.4-5.5, Gemini 3.1 Pro); Polymarket pricing has fallen to 10.5% and is trending down, reflecting real-time market skepticism; re-baselining in July 2026 adds noise/uncertainty rather than clear upward push.
# Gaps / unknowns
- No confirmed, authoritative current top score from arena.ai itself (only third-party aggregators of uncertain reliability).
- No sibling Polymarket threshold prices (1500/1525/1575/1600) retrieved to construct implied distribution.
- Uncertain impact of July 2026 re-baselining on score comparability going forward.
- No Kalshi-direct price returned; anchor relies entirely on Polymarket's identical-question market.
# Calibration anchors
- Polymarket YES price (anchor): 10.5%, declining trend.
- Precedent: sibling 2025 "reach 1500" market resolved NO despite similar optimism pre-launch — historical base rate favors threshold markets resolving NO when large jumps are needed late in the window.