# Current state
No direct arena.ai leaderboard snapshot was retrieved in this research; all coding-Elo figures come from third-party trackers/blogs with inconsistent, sometimes contradictory readings (1462–1582 across "current" claims). The Polymarket-listed price for this same market currently sits at 40.5% YES, down sharply (-28pp) over the past 30 days, suggesting the crowd has become more skeptical the 1560 threshold will be hit by year-end, even as several aggregator sites claim it's already been surpassed.
# Timeline of key events
- 2025-11-18 (confirmed): Gemini 3 Pro becomes first model to break 1500 on LMArena text leaderboard (~1501 Elo); WebDev/coding score ~1487.
- 2025-11-24 (confirmed): Claude Opus 4.5 released; briefly tops WebDev Arena.
- 2026-01 (reported): LMSYS/LMArena rebrands to "Arena."
- 2026-02 (reported): Claude Opus 4.6 leads Coding leaderboard at 1548 (apiyi.com, corroborated by codesota.com).
- 2026-04-23 (reported): Opus 4.7 / 4.7-thinking take top two coding slots; exact score not given.
- 2026-05-24 (reported): Opus 4.7-thinking leads Coding/WebDev at 1567 (propelcode.ai); a separate May 2026 tracker instead cites Opus 4.6 at only 1462 — direct contradiction, unresolved.
- 2026-06-29 (confirmed): TechCrunch reports Arena is now a $100M business (platform credibility/traffic, not score data).
- 2026-07-01/07-12 (reported): Arena performs a score restoration and re-baseline event for at least one model — a methodology change that can shift absolute Elo values discontinuously.
- 2026-07-13 (confirmed): Arena/Google change grading methodology for Android coding evaluation (category refinement).
- 2026-07-17/07-26 (reported): Kimi K3 (open-weight) ships, leads a separate "Frontend Code Arena" at ~1679 — likely a differently-scaled sub-leaderboard, not the resolution source.
- 2026-07-24 (reported): Anthropic launches Claude Opus 5 (agentic coding focus, 1M context).
- 2026-08 (reported, unverified): One tracker (swfte.com) claims Opus 4.8 leads Coding at ~1582, ahead of Opus 4.7 at 1567 — would already exceed 1560, but not corroborated by any primary-source arena.ai screenshot.
- 2026-08 (reported): Grok 5 still not shipped despite rumors; GPT-6 rumored "inside six weeks" as of an August 2026 tracker.
# Event
Will any model on arena.ai's Text Arena "Coding" leaderboard (style control off) reach an Arena Score ≥1560 by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct data was returned in this research pass (kalshi_related found zero matching markets). The only direct market read available is Polymarket on the identical question: **40.5% YES**, down from a 30-day high near 84% (30d change: -28pp; 7d change: -3pp), range 31–84% over 91 days, $90K volume. Treat this as the best available cross-market anchor in lieu of a live Kalshi quote.
# Sub-question answers
1. **Current top coding score/model** — Contested. Reports range from Opus 4.6 at 1462 (one May 2026 tracker) to Opus 4.7-thinking at 1567 (propelcode.ai, May 2026) to Opus 4.8 at ~1582 (swfte.com, Aug 2026). No primary arena.ai screenshot confirms any figure; likely current leader is an Opus 4.7/4.8-class model around 1550–1580, but exact value unverified.
2. **Growth rate (6/12/24mo)** — From ~1487–1501 (Nov 2025, Gemini 3) to reported 1548 (Feb 2026) to 1567 (May 2026) to ~1582 (Aug 2026) implies roughly 5–7 Elo/month over the past ~9 months at the frontier, per claude_news synthesis — though this trend line splices scores from different models/methodology states.
3. **Methodology changes** — Yes: Arena rebranded Jan 2026; performed a "July 1 restoration / July 12 re-baseline" event; refined Coding category filtering (removed non-coding code-like prompts, applied retroactively) — both are documented discontinuities that could shift scores non-organically. [news.lmarena.ai, swfte.com]
4. **2026 frontier releases** — Claude Opus 5 shipped July 24, 2026; GPT-6 rumored ~6 weeks out (as of Aug 2026 reporting); Grok 5 still unreleased (beta rumored May–Jun, API Q3, ~6T param MoE per rumors); Grok 4.6 reportedly imminent; Kimi K3 (open-weight) shipped July 26, 2026, strong on a separate Frontend Code Arena. No confirmed Gemini 3.5/4 in this data.
5. **Gap to 1560** — Ambiguous given source conflict: if leader is truly ~1462, gap is ~100 pts (large, ~19 months to close ≈5.3/mo needed); if leader is ~1548–1567, gap is 0–12 pts (already at/near threshold); if ~1582, threshold already cleared per that source.
6. **Sibling Polymarket/Kalshi threshold markets** — polymarket_related and kalshi_related both returned zero matching sibling markets (no 1500/1520/1540/1580 markets found), so no distributional read-across is available.
# Key facts (high-confidence, factual)
1. [Google blog] Gemini 3 Pro was first to break 1500 overall Elo, Nov 18, 2025.
2. [TechCrunch] Arena platform valued ~$100M business as of June 29, 2026 — actively maintained, resolution source likely to remain online.
3. [news.lmarena.ai] Coding category scoring methodology was refined and retroactively reapplied in 2026, and a re-baseline event occurred July 2026 — scores are not perfectly continuous over time.
4. [Polymarket] Same-question market priced 40.5% YES, having fallen 28pp in 30 days.
# Cross-market signals
- Kalshi related: none found.
- Polymarket: 40.5% YES, strong recent downward momentum (from highs of 84%), suggesting market participants increasingly doubt 1560 will be hit, possibly due to methodology re-baseline lowering effective scores or slower-than-expected progress.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Claude-news synthesis leans bullish (already crossed or very close), citing 1567–1582 August reports.
- Code-execution Monte Carlo model (assuming baseline ~1500) gives a blended ~65-70% central estimate, but is highly sensitive to true starting score and saturation assumptions.
- These optimistic analyst views conflict sharply with Polymarket's declining 40.5% price — a notable disagreement.
# Directional lean per outcome
- **Yes**: Multiple (unverified) trackers claim 1560+ already reached; strong release cadence (Opus 5, rumored GPT-6, Grok 5) into H2 2026; historical ~5-7 Elo/month pace would exceed 1560 easily if sustained.
- **No**: Third-party score reports are wildly inconsistent (1462 vs 1582) and unverified against the actual resolution source; category re-baseline events could suppress scores; Polymarket pricing has fallen sharply to 40.5%, signaling real doubt from an informed trading market; growth may saturate near ceiling.
# Gaps / unknowns
- No direct/current arena.ai screenshot obtained; true current top Coding score unverified.
- Unclear whether the July re-baseline raised or lowered scores.
- No Kalshi-direct price available for primary anchor; relied on Polymarket only.
- No sibling threshold markets to triangulate distribution.
# Calibration anchors
- Polymarket YES price (best direct anchor): 40.5%, trending down.
- Code-execution Monte Carlo central estimate: ~65-70% (sensitive to assumptions, likely overweights aggressive growth continuation).
- Precedent: frontier text/coding Elo has cleared each successive ~50pt band roughly every 3-9 months over 2025-2026, but methodology resets inject real uncertainty.