# Current state
The market resolves YES if any model on arena.ai's official "Coding" leaderboard (style control off) reaches ≥1560 Score by Dec 31, 2026. Third-party aggregator sites (not the official leaderboard) report wildly inconsistent current-leader scores (1487–1679), some already above threshold; Polymarket's 59% price (well below ~100%) suggests the official board has NOT yet confirmed a 1560+ reading, meaning the aggregator claims are likely unreliable, stale, or mis-citing style-control-on numbers.
# Timeline of key events
- 2026-02: "Claude Opus 4.6" reportedly first model past 1500 coding Elo, cited at ~1561 by one aggregator (promptt.dev) — reported, unconfirmed against official board.
- 2026-04-27: buildmvpfast.com snapshot: Opus 4.6 leading Code Arena at 1560 — reported.
- 2026-05-24: propelcode.ai snapshot: claude-opus-4-7-thinking leads at 1567 — reported.
- 2026-06-24: Gemini 3.5 Pro release reportedly slips to July (businessinsider.com) — reported.
- 2026-07-01: "Claude Fable 5" restored after 19-day export-control suspension — reported.
- 2026-07-08/09: Grok 4.5 and GPT-5.6 family (Luna/Terra/Sol) launch — reported.
- 2026-07-13: Google changes Android-coding grading methodology (tech.yahoo.com) — confirmed news event, unclear leaderboard impact.
- 2026-07-16: Kimi K3 (Moonshot, open-weight) launches — reported.
- 2026-07-19: Alibaba previews Qwen3.8, claims #2 behind Claude Fable 5 — reported.
- 2026-07-27: benchlm.ai snapshot: claude-fable-5 leads at 1553 — reported (lower than some June/July figures, inconsistent).
- 2026-08 (swfte.com): Opus 4.8 ~1582 overall lead cited, but Kimi K3 separately claimed #1 on coding at 1679 — reported, internally inconsistent across sources.
# Event
Will any AI model reach a Coding Arena Score of 1560+ on arena.ai's Text Arena "Coding" leaderboard (style control off) by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct price returned (only Polymarket data available as primary anchor here). **Polymarket price: 59% YES**, down 9.5% over 7 days, up 12% over 30 days; range 31–84% over 91 days; volume $85K. High volatility suggests market is reacting to conflicting/unclear leaderboard news rather than converging.
# Sub-question answers
1. **Current top score/model** — Disputed. Aggregator sites give 1487 (code_execution baseline), 1553 (benchlm.ai), 1560 (buildmvpfast.com), 1567 (propelcode.ai), 1582/1679 (swfte.com) — no confirmed single official reading; likely leader is a Claude Opus/Fable variant, possibly Kimi K3 on some readings.
2. **Rate of rise** — Overall Text Arena rose ~122 Elo in 2024→2025, then ~42 Elo in 5 months into 2026 (~100/yr annualized per swfte.com); coding-specific trend is lumpier, driven by discrete flagship releases (Opus 4.6→4.7→4.8 jumps of ~7-15 pts each per reported snapshots).
3. **2026 frontier releases** — Gemini 3.5 Pro (slipped to July), GPT-5.6 family (Luna/Terra/Sol, July), Grok 4.5 (July), Claude Fable 5/Opus 4.8, Kimi K3 (Moonshot open-weight), Qwen3.8 — active release cadence continuing through Q3 2026, several claimed to already threaten/exceed 1560.
4. **Score compression/rebaselining** — A LMArena rebrand/methodology update (Style Control refinement) reportedly shifted Elo distributions ±20-40 points without real quality change (agileleadershipdayindia.org) — adds material uncertainty to whether "1560" is comparable across time; no hard rebaselining cap identified.
5. **Gap to close** — Per code_execution baseline (~1487), gap = ~73 points over ~15 months (≈4.9 pts/mo breakeven). If aggregator claims of 1560+ already being reached are accurate, gap = 0.
6. **Polymarket/sibling markets** — Only this single market found (no 1500/1520/1540 sibling markets on Polymarket); no additional distributional read available.
# Key facts (high-confidence, factual)
1. [Wikipedia] LMArena (formerly Chatbot Arena) is a live, continuously-updated crowdsourced Elo leaderboard; scores are not fixed/absolute across time.
2. [Polymarket] Current YES price 59%, volatile (31-84% range over 91 days).
3. [arena.ai/blog/leaderboard-changelog, via claude_news] New models (Kimi K2.7-code, Opus 4.8/4.8-thinking, Mistral Medium 3.5) are continually added to the official Code leaderboard, confirming active leaderboard updates in 2026.
4. [gdelt] Confirmed real-world events: Gemini 3.5 Pro delay to July 2026; Grok 4.5 debut; Google changed Android-coding grading methodology (July 2026).
# Cross-market signals
- Kalshi related: no direct sibling market found (only irrelevant "AI model" keyword match).
- Polymarket: 59% YES, no sibling threshold markets found.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Multiple SEO/aggregator blogs (benchlm.ai, swfte.com, propelcode.ai, buildmvpfast.com) claim the 1560 threshold has already been crossed multiple times in 2026, but these are non-primary sources with mutually inconsistent numbers (1553-1679 range) and are explicitly flagged by claude_news itself as unreliable/directional only.
- code_execution's structural growth model (independent of these claims) puts current baseline at ~1487, implying ~73-point gap and roughly coin-flip (50-55%) probability if trend continues at 5 pts/month, consistent with Polymarket's 59%.
# Directional lean per outcome
- **Yes**: Rapid 2026 release cadence (Opus 4.6→4.8, GPT-5.6, Grok 4.5, Kimi K3, Gemini 3.5) with several unofficial readings already at/above 1560; strong historical rate of coding-score growth; long runway to Dec 2026.
- **No**: Official leaderboard reading not confirmed to have crossed 1560 (Polymarket at 59%, not near 100%, argues against threshold already being met); methodology/rebaselining changes could suppress comparability; aggregator numbers are unreliable and contradictory.
# Gaps / unknowns
- No direct read of the official arena.ai/leaderboard/text/coding-no-style-control page was obtained — all "current score" data is third-hand/aggregator-derived and conflicting.
- Unclear whether recent Style Control methodology changes affect the specific "no style control" Coding tab used for resolution.
- No Kalshi-direct price was returned; Polymarket used as sole cross-market anchor.
# Calibration anchors
- Polymarket YES price: **59%** (primary anchor, but notably below near-certainty despite aggregator claims of threshold already crossed — a key tension).
- code_execution structural model: breakeven ≈4.9 pts/month; at plausible 5 pts/month trend, P≈54-55%, aligning closely with Polymarket price.
- Precedent: coding Elo crossed 1500 for first time only in ~Feb 2026 per one source; 1560 would represent a further ~60-point jump requiring 1-2 more flagship-tier releases — plausible but not guaranteed within remaining ~15 months.