# Current state
Polymarket prices this event at 37.5% YES, down sharply from an 81.5% high and a -26.5% 30-day move — signaling real market skepticism about crossing 1560 on the specific "no-style-control Coding" leaderboard, despite third-party blogs citing already-above-1560 scores (which likely reference different leaderboard variants). The resolution source (arena.ai/leaderboard/text/coding-no-style-control) has not been directly queried in this research; all cited scores come from secondary trackers with methodology ambiguity.
# Timeline of key events
- 2026-05-24 (reported): "Coding" leaderboard snapshot shows claude-opus-4-7-thinking leading at 1567 (propelcode.ai) — but this cites the "Code Arena WebDev" leaderboard, not necessarily the resolution-specified no-style-control Coding tab.
- 2026-06-09 (reported): Anthropic launches Claude Fable 5 / Mythos 5 (morphllm.com, benchlm.ai).
- Mid-2026 (reported): Claude Opus 4.7 cited as #1 on "Hard Prompts and Coding" with overall Elo ~1420 — a different/lower scale than the 1567-1582 figures (toolcenter.ai), indicating scale inconsistency across trackers.
- 2026-07-01/07-12 (reported): Arena.ai undergoes a service restoration and re-baselines scores to count only post-restoration votes — a methodology reset that breaks score continuity (localaimaster.com).
- 2026-07-19 (reported): Alibaba previews Qwen3.8, claiming second place behind Claude Fable 5 (siliconangle.com/GDELT).
- 2026-07-24 (reported): Claude Opus 5 reaches GA (morphllm.com).
- 2026-08-04 (reported): "Fable 5 Laps Field on MirrorCode" — GPT-5.5 coding score reportedly collapses on a benchmark redesign (techtimes.com/GDELT), underscoring volatility/benchmark-design sensitivity of coding scores.
- 2026-08 (reported): swfte.com tracker claims Claude Opus 4.8 leads plain Coding Arena at ~1582, ahead of Opus 4.7 at 1567 — unverified against the official style-control-off resolution page.
- 2026-08 (reported): A separate "Fullstack Code Arena" (distinct sub-benchmark) shows Claude Opus 5 Max at 1699 and GPT-5.6 Sol at ~1638 — NOT the plain Coding category used for resolution.
- Ongoing 2026-08/09: Frontier release cadence continues (Gemini 3.7 Flash, GPT-5.6 tiers, Grok 4.6, DeepSeek V4-Pro) roughly monthly (GDELT/heise.de/memeburn.com).
# Event
Will any model reach ≥1560 on the arena.ai Text Arena "Coding" leaderboard (style control OFF) by Dec 31, 2026?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No Kalshi-direct price returned in this research batch (kalshi_related found 0 matches). Treat Polymarket as the primary cross-market anchor: **37.5% YES**, down from an 81.5% peak, -26.5% over 30 days, +2% over 7 days, on $93K volume — a meaningful downward repricing suggesting the market believes the threshold has NOT yet been cleanly crossed on the specific resolution page.
# Sub-question answers
1. **Current top score / model** — Conflicting: third-party trackers claim 1567–1582 (Claude Opus 4.7/4.8) on "Coding," but these appear to reference WebDev/Hard-Prompts variants, not confirmed to be the exact "coding-no-style-control" page. No direct read of the official resolution URL was obtained.
2. **Trend rate of increase** — Not cleanly quantified; Arena's own blog says top-5 mean score rose from ~1000 (May 2023) to ~1500 (overall Text Arena, Aug 2026), implying long-run gradual drift plus release-driven jumps of ~10-40 pts (claude_news, code_execution model).
3. **Gap to 1560 and time needed** — If current top score is genuinely ~1567+, gap is already closed; if it's closer to 1490-1520 (per Sophon style-control tracker at 1550, or toolcenter's 1420 overall), 10-90 points remain, closable in 1-6 months at typical jump sizes per Monte Carlo model.
4. **Frontier releases expected before Dec 2026** — Confirmed pattern of near-monthly frontier releases (Opus 5, GPT-5.6 tiers, Gemini 3.7, Grok 4.6, DeepSeek V4-Pro, Qwen3.8) through Aug 2026; cadence strongly supports continued step-jumps (GDELT, claude_news).
5. **Methodology changes** — Confirmed: Arena re-baselined scores July 12, 2026 after a July 1 restoration, resetting vote-counting — this breaks trend continuity and adds uncertainty to any "current score" reading (localaimaster.com).
6. **Polymarket price / related markets** — 37.5% YES on this exact market; no related Polymarket or Kalshi markets found for adjacent thresholds/categories to triangulate a distribution.
# Key facts (high-confidence, factual)
1. [polymarket_direct] Current Polymarket YES price: 37.5%, down from 81.5% high, -26.5% over 30 days.
2. [localaimaster.com] Arena re-baselined leaderboard scoring on 2026-07-12.
3. [morphllm.com] Claude Opus 5 reached GA 2026-07-24; Opus 4.8 released 2026-05-28.
4. [GDELT/techtimes.com] Benchmark redesign caused a reported GPT-5.5 coding score "collapse" (2026-08-04), showing scores are sensitive to benchmark/methodology changes, not just model capability.
5. [Wikipedia] Arena (formerly LMArena/Chatbot Arena) has documented methodological limitations noted in independent research.
# Cross-market signals
- Kalshi related: none found.
- Polymarket: 37.5% YES, high volatility (31%-81.5% range over 91 days), recent downtrend dominant despite small 7-day uptick.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- claude_news synthesis argues trajectory "strongly favors crossing 1560... may have already occurred," but this leans on trackers not confirmed to match the exact resolution page/style-control setting.
- code_execution Monte Carlo (not grounded in confirmed current score) estimates 70-96% probability range, centered ~85%, but is highly sensitive to unverified starting score assumption (S0).
# Directional lean per outcome
- **Yes**: Frequent frontier releases (monthly cadence), historical jump sizes (10-40 pts), long-run upward drift, and some trackers claiming score already >1560 all support Yes.
- **No**: Polymarket's sharp 30-day decline to 37.5% (from 81.5%) suggests informed traders see the *specific* no-style-control resolution score still below 1560; re-baselining/methodology resets add score-continuity risk; conflicting leaderboard variants (style-control 1550, overall Hard-Prompts 1420) suggest true "coding-no-style-control" score may be lower than optimistic trackers claim.
# Gaps / unknowns
- No direct read of arena.ai/leaderboard/text/coding-no-style-control obtained — the single most important missing data point.
- No Kalshi-direct price captured for this ticker.
- Unclear why Polymarket price fell so much if scores are truly already >1560 — possible resolution ambiguity or trader skepticism about tracker accuracy.
# Calibration anchors
- Polymarket YES price (anchor): 37.5%, recent 30-day decline of -26.5pp.
- Monte Carlo base-rate model (unverified inputs): ~70-90% range.
- Precedent: rapid, frequent frontier-model releases historically produce periodic 10-40 pt Arena jumps, but methodology resets (as seen July 2026) can offset apparent progress.