# Current state
The event resolves off LiveBench.ai's "Coding" category leaderboard score checked at 2026-08-31 12:00 PM ET; whichever company owns the top-scoring model wins. As of mid-August 2026, no single source confirms who currently leads the *specific* LiveBench Coding sub-category — Anthropic's Claude Fable 5/Opus 5 lead LiveBench's *overall* score and several third-party coding aggregators (BenchLM), while Google's Gemini 3 Pro Preview leads the related-but-distinct LiveCodeBench, and OpenAI's GPT-5.5/5.6 leads on SWE-bench Verified (independent eval). The market itself (current YES 90%) is pricing Anthropic as heavy favorite.
# Timeline of key events
- 2025-11 to 2026-01: Robinhood/Kalshi market on "top LiveBench Coding Average" priced OpenAI at 99¢ entering Jan 2026 — OpenAI held the specific Coding leaderboard lead at that time (reported, robinhood.com).
- 2026-06-20/23: Small/open models (tiny model, Sakana AI "Fugu") claim to beat Claude Fable 5 on coding benchmarks (reported, propakistani.pk/moneycontrol.com) — noise, not leaderboard-determinative.
- 2026-06-25: LiveBench does its routine monthly contamination-control question rotation; Claude Fable 5 leads overall LiveBench snapshot at 83.0%, GPT-5.6 Sol 81.1%, GPT-5.5 80.2% (reported, benchlm.ai).
- 2026-06-29/30: Reports of Gemini 3.5 Pro "cleared for July launch," Fable 5 "nearing return," GPT-5.6 "still locked" (rumored, techtimes.com).
- 2026-07-13: Google changes Android-coding grading methodology (reported, Yahoo Tech); GLM-5.2 captures 40% developer tokens (reported, techtimes.com).
- 2026-07-16: Bloomberg reports Gemini 3.5 Pro missed its third internal deadline, months behind schedule, coding capability below Google's internal bar (reported).
- 2026-07-18: Kimi K3 found to be an illegal Claude distillation (reported, propakistani.pk) — reputational noise for Anthropic ecosystem.
- 2026-07-24/25: Anthropic launches Claude Opus 5 — large SWE-bench Pro gains (69.2%→79.2%), doubles Opus 4.8 on Frontier-Bench agentic coding, same price as Opus 4.8 (confirmed via multiple outlets: arynews.tv, iclarified.com, codersera.com).
- 2026-08-05/14/15: Google ships Gemini 3.6 Flash and Gemini 3.7 Flash (mid-tier, not flagship Pro replacement) (reported, thenextweb.com/digitaltrends.com/heise.de).
- 2026-08-12: LiveCodeBench (distinct benchmark) shows Gemini 3 Pro Preview #1 at 91.7%, ahead of Gemini 3 Flash Preview and DeepSeek V3.2 (reported, pricepertoken.com).
- 2026-08 (BenchLM aggregate): Claude Mythos 5 leads composite coding ranking (81.1), ahead of Claude Fable 5 (80.8) and GPT-5.6 Sol (78.7) (reported, benchlm.ai/coding).
# Event
Will Anthropic (vs Google/OpenAI/xAI/Other) own the #1-scoring model in LiveBench.ai's "Coding" category as checked 2026-08-31 12:00 PM ET?
# Outcomes to forecast
Yes (Anthropic tops LiveBench Coding) / No (any other company tops it)
# Kalshi market anchor
Cross-listed market data (same ticker) shows current price **90.0%** YES for Anthropic, up +3pp over 7 days and +40pp over 30 days (price range 50%–93.5% over 19 days of data). Volume modest (~$16K total). Strong recent momentum toward Yes, likely driven by the Opus 5 launch (2026-07-24/25) and subsequent benchmark chatter.
# Sub-question answers
1. **Current LiveBench Coding leader / gap** — Not definitively confirmed for the exact "Coding" sub-category as of mid-Aug 2026; Anthropic leads LiveBench *overall* (Fable 5, 83.0%) and BenchLM's composite coding score, but a Jan-2026 Coding-specific market had OpenAI at 99%, and LiveCodeBench (a different benchmark) currently favors Google's Gemini 3 Pro. No source gives a clean current Coding-category score gap.
2. **Frequency of leadership change / base rate** — No direct historical cadence data found; illustrative code_execution model shows that under naive monthly-flip assumptions (p=0.2–0.5), persistence over the ~1-year horizon would be <15%, well below market's 90% — implying leadership is either much "stickier" than a memoryless model or market is pricing genuine Anthropic-specific durability.
3. **Is LiveBench still active in 2026?** — Yes; confirmed to have done a routine monthly contamination-control refresh on 2026-06-25 (pinggy.io), indicating active maintenance into Q3 2026.
4. **New Anthropic releases before Aug 2026** — Claude Opus 5 launched 2026-07-24/25 with major coding gains (SWE-bench Pro 69.2%→79.2%, Frontier-Bench agentic coding more than doubled vs Opus 4.8), same price tier (confirmed, multiple outlets). Claude Fable 5/Mythos 5 also referenced as top-tier Anthropic coding models in mid-2026.
5. **Competing releases (Google/OpenAI/xAI/DeepSeek)** — Google shipped Gemini 3.6/3.7 Flash (mid-tier) but flagship Gemini 3.5 Pro is delayed (Bloomberg, 2026-07-16, "months behind schedule"); OpenAI's GPT-5.5/5.6 Sol lead SWE-bench Verified independently; xAI shipped Grok 4.5 and Grok Build (agentic) in June-July 2026; DeepSeek V3.2/V4-Pro and Qwen3.8-Max also active but not leading closed-model benchmarks.
6. **Polymarket sibling market prices** — Only this Anthropic-outcome market data was retrieved directly (90%); no confirmed Google/OpenAI/xAI/Other sibling prices found (polymarket_related search returned 0 matches). A hypothetical illustrative de-vig (not live data) suggested Anthropic ~42%, Google ~27%, OpenAI ~20%, xAI ~7%, Other ~4% — but this is explicitly labeled illustrative, not live.
# Key facts (high-confidence, factual)
1. [polymarket_direct] Current YES price for this exact market: 90%, +40pp over 30 days.
2. [pinggy.io] LiveBench actively refreshed questions 2026-06-25, confirming ongoing maintenance.
3. [multiple] Claude Opus 5 launched 2026-07-24/25 with substantial coding benchmark gains at flat pricing.
4. [Bloomberg via felloai.com/techtimes] Google's Gemini 3.5 Pro flagship delayed past three deadlines as of mid-July 2026.
5. [benchlm.ai] BenchLM's composite coding ranking (Aug 2026) has two Claude variants (Mythos 5, Fable 5) in top 2 spots.
6. [pricepertoken.com] Gemini 3 Pro Preview leads LiveCodeBench (different benchmark) as of 2026-08-12.
# Cross-market signals
- Kalshi related: Anthropic IPO-first market at 93% (confirms strong overall market confidence in Anthropic's momentum/position, but unrelated to coding benchmarks specifically).
- Polymarket: No sibling outcome prices confirmed live; only Anthropic-outcome price found (90%).
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Multiple benchmark aggregators (BenchLM, felloai.com) argue Anthropic's Opus 5/Fable 5/Mythos 5 lineage has "persistent" coding leadership through mid-2026, aided by Google's Gemini 3.5 Pro delay removing the most likely near-term challenger.
- Others note "best" model is highly benchmark-dependent (SWE-bench vs LiveCodeBench vs LiveBench Coding vs Arena WebDev), with OpenAI and Google each winning on at least one major coding metric.
# Directional lean per outcome
- **Yes (Anthropic)**: Opus 5 launch strength, LiveBench-overall lead, BenchLM composite lead, Gemini 3.5 Pro delay removing top rival, strong and rising market price (90%, +40pp/30d).
- **No (other)**: Historical Jan-2026 data had OpenAI leading LiveBench Coding specifically; Gemini leads LiveCodeBench; GPT-5.5/5.6 leads SWE-bench Verified; benchmark leadership in this space has shown volatility across releases (new GPT-5.x/Gemini-3.x variants could flip the specific Coding sub-score before Aug 31).
# Gaps / unknowns
- No confirmed direct read of the actual LiveBench.ai "Coding" category score/ranking as of the current date — all evidence is inferential from overall LiveBench score or other benchmarks.
- No live Polymarket sibling-market prices for Google/OpenAI/xAI/Other confirmed (illustrative figures only).
- Uncertain whether OpenAI or Google will ship a coding-focused frontier update before 2026-08-31 that could flip the specific Coding leaderboard.
# Calibration anchors
- Kalshi/Polymarket current YES price (anchor): 90%.
- Jan-2026 sibling market precedent: OpenAI priced at 99% for LiveBench Coding Average lead — shows category leadership can be decisively one company at a given snapshot, and has flipped since (from OpenAI to contested/Anthropic-leaning).