← Back to scans

Will Anthropic have the best AI model on LiveBench (Coding) at the end of August 2026?

0x272c4523eec6df3d46f668c0e10aa3bf91059074016a66d1e042a4c2118533d9 · Science and Technology · 2026-08-16
85%
Agent
90%
Market Price
-5.0%
Edge
59%
Confidence
Volume: 15,971
Spread: 2.0c
Days to resolution: 15
Markets in event: 32
Final Rationale
The Kalshi anchor (90% YES, +40pp/30d) plus a very short remaining horizon (~2 weeks) makes Anthropic the clear favorite: Claude Opus 5's late-July coding gains landed, Anthropic leads LiveBench overall and BenchLM's composite coding ranking, and the most credible challenger (Gemini 3.5 Pro) is reportedly months behind schedule with Google only shipping mid-tier Flash models in August. However, the critique lands on real weaknesses: no source directly confirms the LiveBench *Coding* sub-category ranking, OpenAI held that exact category at 99% as recently as Jan-2026, the BenchLM gap is razor-thin (81.1 vs 78.7), and LiveBench's monthly contamination-control question rotation injects exogenous variance right before the Aug-31 snapshot. The naive monthly-flip base rate is not the right model over a 2-week window — leadership is empirically sticky between frontier releases — so it justifies only a modest, not dramatic, discount. I therefore shade a few points below both the market and Forecast 1, landing at 0.85, which also implicitly leaves room for tail flips from OpenAI's pending GPT-5.x update or an unexpected xAI/open-weight benchmark spike.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 15$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related claude_news claude_news gdelt_news kalshi_related code_execution wikipedia
Sub-questions (Fermi decomposition)
  1. Which company's model currently holds the highest Coding score on the LiveBench.ai leaderboard, and what is the score gap to the #2 model?
  2. How often has LiveBench Coding leadership changed hands over the past 12-18 months, and what is the base rate of the current leader still leading ~1-2 months later?
  3. Does LiveBench still actively update its leaderboard in 2026, and how quickly are new frontier models (Claude, Gemini, GPT, Grok) added?
  4. What new or upcoming Anthropic Claude models (e.g., Claude Opus/Sonnet 4.x or 5) are released or expected before Aug 31, 2026, and how do they benchmark on coding?
  5. What competing frontier releases from Google (Gemini 3.x), OpenAI (GPT-5.x), xAI (Grok 5), and DeepSeek/Qwen are expected in the same window?
  6. What are the current Polymarket prices for the sibling markets (Google, OpenAI, xAI, Other) in this same 'best AI model on LiveBench Coding' event group, and what do they imply after de-vigging?
Planner reasoning
This is a Polymarket question about which company tops the LiveBench Coding leaderboard on Aug 31, 2026, so the market's own price is the primary anchor, plus the sibling markets for Google/OpenAI/xAI in the same event group give an implied full distribution. Key empirical drivers are: who currently leads LiveBench Coding, by what margin, how frequently the leaderboard is updated/leadership changes, and what frontier model releases (Claude, Gemini, GPT) are expected before end of August 2026.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.5s 1 ## This Market's Polymarket Data **Will Anthropic have the best AI model on LiveBench (Coding) at the end of August 2026?** - Current price (probability): 90.00% - 7-day price change: +3.00% - 30-day price change: +40.00% - Total volume: $15,971 (USD notional) - Price range: 50.00% - 93.50% - Data
polymarket_related OK 2.5s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'LiveBench': 0 markets | keyword 'best AI model': 0 markets | keyword 'Coding leaderboard': 0 markets | keyword 'Anthropic': 0 markets | keyword 'Gemini': 0 markets
claude_news OK 32.4s 8 Based on my research, here are key findings on LiveBench.ai Coding leaderboard standings relevant to this question: - **LiveBench is actively updated with monthly question rotations for contamination control.** LiveBench refreshed its question set on June 25, 2026, part of its normal contamination
claude_news OK 27.3s 11 Based on research, here are key findings on the AI coding leaderboard landscape approaching August 31, 2026: - **Anthropic currently leads several coding leaderboards as of mid-August 2026**: the AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, base
gdelt_news OK 150.2s 20 GDELT: 20 articles across 3 queries (lookback=60d). 'LiveBench coding leaderboard model': error GDELT rate-limited after retries (429) | 'Claude coding benchmark best model': 10 hits | 'Gemini GPT coding benchmark leader 2026': 10 hits
kalshi_related OK 2.4s 2 2 related markets / summaries. keyword 'best AI model': ok | keyword 'LiveBench': no matches | keyword 'Anthropic': ok
code_execution OK 23.9s 0 **Note:** No exact live Polymarket price snapshot was provided in the prompt, so the de-vig calculation below uses a representative/illustrative set of market prices for the five outcomes (Anthropic, Google, OpenAI, xAI, Other); replace with live quotes for a precise real-time read. - **Raw illustr
wikipedia OK 0.1s 3 Fetched 3 Wikipedia entries (0 missing pages).
3. Evidence Brief Sonnet · 8216 chars
# Current state The event resolves off LiveBench.ai's "Coding" category leaderboard score checked at 2026-08-31 12:00 PM ET; whichever company owns the top-scoring model wins. As of mid-August 2026, no single source confirms who currently leads the *specific* LiveBench Coding sub-category — Anthropic's Claude Fable 5/Opus 5 lead LiveBench's *overall* score and several third-party coding aggregators (BenchLM), while Google's Gemini 3 Pro Preview leads the related-but-distinct LiveCodeBench, and OpenAI's GPT-5.5/5.6 leads on SWE-bench Verified (independent eval). The market itself (current YES 90%) is pricing Anthropic as heavy favorite. # Timeline of key events - 2025-11 to 2026-01: Robinhood/Kalshi market on "top LiveBench Coding Average" priced OpenAI at 99¢ entering Jan 2026 — OpenAI held the specific Coding leaderboard lead at that time (reported, robinhood.com). - 2026-06-20/23: Small/open models (tiny model, Sakana AI "Fugu") claim to beat Claude Fable 5 on coding benchmarks (reported, propakistani.pk/moneycontrol.com) — noise, not leaderboard-determinative. - 2026-06-25: LiveBench does its routine monthly contamination-control question rotation; Claude Fable 5 leads overall LiveBench snapshot at 83.0%, GPT-5.6 Sol 81.1%, GPT-5.5 80.2% (reported, benchlm.ai). - 2026-06-29/30: Reports of Gemini 3.5 Pro "cleared for July launch," Fable 5 "nearing return," GPT-5.6 "still locked" (rumored, techtimes.com). - 2026-07-13: Google changes Android-coding grading methodology (reported, Yahoo Tech); GLM-5.2 captures 40% developer tokens (reported, techtimes.com). - 2026-07-16: Bloomberg reports Gemini 3.5 Pro missed its third internal deadline, months behind schedule, coding capability below Google's internal bar (reported). - 2026-07-18: Kimi K3 found to be an illegal Claude distillation (reported, propakistani.pk) — reputational noise for Anthropic ecosystem. - 2026-07-24/25: Anthropic launches Claude Opus 5 — large SWE-bench Pro gains (69.2%→79.2%), doubles Opus 4.8 on Frontier-Bench agentic coding, same price as Opus 4.8 (confirmed via multiple outlets: arynews.tv, iclarified.com, codersera.com). - 2026-08-05/14/15: Google ships Gemini 3.6 Flash and Gemini 3.7 Flash (mid-tier, not flagship Pro replacement) (reported, thenextweb.com/digitaltrends.com/heise.de). - 2026-08-12: LiveCodeBench (distinct benchmark) shows Gemini 3 Pro Preview #1 at 91.7%, ahead of Gemini 3 Flash Preview and DeepSeek V3.2 (reported, pricepertoken.com). - 2026-08 (BenchLM aggregate): Claude Mythos 5 leads composite coding ranking (81.1), ahead of Claude Fable 5 (80.8) and GPT-5.6 Sol (78.7) (reported, benchlm.ai/coding). # Event Will Anthropic (vs Google/OpenAI/xAI/Other) own the #1-scoring model in LiveBench.ai's "Coding" category as checked 2026-08-31 12:00 PM ET? # Outcomes to forecast Yes (Anthropic tops LiveBench Coding) / No (any other company tops it) # Kalshi market anchor Cross-listed market data (same ticker) shows current price **90.0%** YES for Anthropic, up +3pp over 7 days and +40pp over 30 days (price range 50%–93.5% over 19 days of data). Volume modest (~$16K total). Strong recent momentum toward Yes, likely driven by the Opus 5 launch (2026-07-24/25) and subsequent benchmark chatter. # Sub-question answers 1. **Current LiveBench Coding leader / gap** — Not definitively confirmed for the exact "Coding" sub-category as of mid-Aug 2026; Anthropic leads LiveBench *overall* (Fable 5, 83.0%) and BenchLM's composite coding score, but a Jan-2026 Coding-specific market had OpenAI at 99%, and LiveCodeBench (a different benchmark) currently favors Google's Gemini 3 Pro. No source gives a clean current Coding-category score gap. 2. **Frequency of leadership change / base rate** — No direct historical cadence data found; illustrative code_execution model shows that under naive monthly-flip assumptions (p=0.2–0.5), persistence over the ~1-year horizon would be <15%, well below market's 90% — implying leadership is either much "stickier" than a memoryless model or market is pricing genuine Anthropic-specific durability. 3. **Is LiveBench still active in 2026?** — Yes; confirmed to have done a routine monthly contamination-control refresh on 2026-06-25 (pinggy.io), indicating active maintenance into Q3 2026. 4. **New Anthropic releases before Aug 2026** — Claude Opus 5 launched 2026-07-24/25 with major coding gains (SWE-bench Pro 69.2%→79.2%, Frontier-Bench agentic coding more than doubled vs Opus 4.8), same price tier (confirmed, multiple outlets). Claude Fable 5/Mythos 5 also referenced as top-tier Anthropic coding models in mid-2026. 5. **Competing releases (Google/OpenAI/xAI/DeepSeek)** — Google shipped Gemini 3.6/3.7 Flash (mid-tier) but flagship Gemini 3.5 Pro is delayed (Bloomberg, 2026-07-16, "months behind schedule"); OpenAI's GPT-5.5/5.6 Sol lead SWE-bench Verified independently; xAI shipped Grok 4.5 and Grok Build (agentic) in June-July 2026; DeepSeek V3.2/V4-Pro and Qwen3.8-Max also active but not leading closed-model benchmarks. 6. **Polymarket sibling market prices** — Only this Anthropic-outcome market data was retrieved directly (90%); no confirmed Google/OpenAI/xAI/Other sibling prices found (polymarket_related search returned 0 matches). A hypothetical illustrative de-vig (not live data) suggested Anthropic ~42%, Google ~27%, OpenAI ~20%, xAI ~7%, Other ~4% — but this is explicitly labeled illustrative, not live. # Key facts (high-confidence, factual) 1. [polymarket_direct] Current YES price for this exact market: 90%, +40pp over 30 days. 2. [pinggy.io] LiveBench actively refreshed questions 2026-06-25, confirming ongoing maintenance. 3. [multiple] Claude Opus 5 launched 2026-07-24/25 with substantial coding benchmark gains at flat pricing. 4. [Bloomberg via felloai.com/techtimes] Google's Gemini 3.5 Pro flagship delayed past three deadlines as of mid-July 2026. 5. [benchlm.ai] BenchLM's composite coding ranking (Aug 2026) has two Claude variants (Mythos 5, Fable 5) in top 2 spots. 6. [pricepertoken.com] Gemini 3 Pro Preview leads LiveCodeBench (different benchmark) as of 2026-08-12. # Cross-market signals - Kalshi related: Anthropic IPO-first market at 93% (confirms strong overall market confidence in Anthropic's momentum/position, but unrelated to coding benchmarks specifically). - Polymarket: No sibling outcome prices confirmed live; only Anthropic-outcome price found (90%). - Sportsbook implied: N/A. # Analyst opinions and speculation - Multiple benchmark aggregators (BenchLM, felloai.com) argue Anthropic's Opus 5/Fable 5/Mythos 5 lineage has "persistent" coding leadership through mid-2026, aided by Google's Gemini 3.5 Pro delay removing the most likely near-term challenger. - Others note "best" model is highly benchmark-dependent (SWE-bench vs LiveCodeBench vs LiveBench Coding vs Arena WebDev), with OpenAI and Google each winning on at least one major coding metric. # Directional lean per outcome - **Yes (Anthropic)**: Opus 5 launch strength, LiveBench-overall lead, BenchLM composite lead, Gemini 3.5 Pro delay removing top rival, strong and rising market price (90%, +40pp/30d). - **No (other)**: Historical Jan-2026 data had OpenAI leading LiveBench Coding specifically; Gemini leads LiveCodeBench; GPT-5.5/5.6 leads SWE-bench Verified; benchmark leadership in this space has shown volatility across releases (new GPT-5.x/Gemini-3.x variants could flip the specific Coding sub-score before Aug 31). # Gaps / unknowns - No confirmed direct read of the actual LiveBench.ai "Coding" category score/ranking as of the current date — all evidence is inferential from overall LiveBench score or other benchmarks. - No live Polymarket sibling-market prices for Google/OpenAI/xAI/Other confirmed (illustrative figures only). - Uncertain whether OpenAI or Google will ship a coding-focused frontier update before 2026-08-31 that could flip the specific Coding leaderboard. # Calibration anchors - Kalshi/Polymarket current YES price (anchor): 90%. - Jan-2026 sibling market precedent: OpenAI priced at 99% for LiveBench Coding Average lead — shows category leadership can be decisively one company at a given snapshot, and has flipped since (from OpenAI to contested/Anthropic-leaning).
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.60
Yes 89%
No 11%
The market anchor is 90% YES with strong momentum (+40pp over 30 days), consistent with Claude Opus 5's late-July launch (SWE-bench Pro 69.2%→79.2%, doubled agentic coding scores) and Anthropic's lead on LiveBench overall (Fable 5 at 83.0%) plus BenchLM's composite coding ranking. The most credible challenger, Google's Gemini 3.5 Pro, is reportedly months behind schedule per Bloomberg, and Google's August releases were mid-tier Flash models, leaving only two weeks for a surprise frontier drop from OpenAI or Google to flip the specific Coding sub-score. Countervailing: no source directly confirms the current LiveBench Coding sub-category ranking (Anthropic's lead is inferred from overall score and other aggregators), Gemini 3 Pro leads the distinct LiveCodeBench, and OpenAI's GPT-5.6 leads SWE-bench Verified — so a modest discount from the anchor is warranted. With a short horizon and no identified imminent competitor launch, I stay essentially at the market price.
gpt-5.5
0.58
Yes 82%
No 18%
The Kalshi anchor is a strong 90% YES and has moved sharply toward Anthropic, likely reflecting Claude Opus 5's late-July coding gains and Anthropic's apparent LiveBench/BenchLM momentum. The current-leader evidence points toward Yes because Anthropic leads LiveBench overall and BenchLM composite coding, though it is weakened by the lack of a direct LiveBench Coding subcategory read and by Google/OpenAI leading other coding benchmarks. The release-cycle evidence also points toward Yes because Opus 5 delivered large coding improvements while Google's likely flagship challenger is reportedly delayed, but historical OpenAI leadership in this exact category and benchmark volatility leave meaningful flip risk. I therefore stay anchored high but shade below market to account for the unresolved exact leaderboard gap and possible late OpenAI/Google updates before the August 31 check.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts anchor heavily on the 90% Kalshi price without adequately weighting sub-question #1's key finding: "No source gives a clean current Coding-category score gap," and the Jan-2026 precedent where OpenAI held 99% on this *exact* category just months earlier — this suggests the specific-category leaderboard is more volatile/flip-prone than the market's overall-score-driven momentum implies, warranting a larger discount than 8-11pp. 2. Neither forecast grapples with the base-rate math explicitly flagged in the brief: "under naive monthly-flip assumptions (p=0.2–0.5), persistence over the ~1-year horizon would be <15%, well below market's 90%" — they should have either explained why leadership is "stickier" than this model or discounted further; instead both simply defer to the anchor. 3. Both forecasts under-weight the possibility that OpenAI (which leads SWE-bench Verified and held Coding-specific leadership as recently as Jan 2026) could ship a coding-focused update in the final weeks before Aug 31 — the brief explicitly flags "GPT-5.6 still locked" rumors suggesting a pending release, and neither analyst separately assesses OpenAI's flip probability vs. Google's (they treat "No" as a monolithic residual dominated by Google's delay). 4. The BenchLM composite score gap cited as support for Anthropic is razor-thin (Mythos 5 81.1 vs Fable 5 80.8 vs GPT-5.6 Sol 78.7) — this is not a commanding lead and could easily flip with routine LiveBench's monthly contamination-control question rotation (confirmed to occur, e.g., 2026-06-25), a mechanism neither forecaster discusses as a source of exogenous variance right before the Aug 31 checkpoint. 5. Neither forecast addresses the "Other" category (xAI's Grok 4.5/Grok Build, DeepSeek V4-Pro, Qwen3.8-Max) as a tail risk — while unlikely, both treat No as effectively synonymous with Google/OpenAI, ignoring that an open-weight or xAI model could top a specific automated benchmark unexpectedly, a scenario the illustrative de-vig (Anthropic ~42%, Google ~27%, OpenAI ~20%, xAI ~7%, Other ~4%) at least allocates probability to but the final forecasts do not explicitly reconcile with. 6. Both forecasts collapse to within 7pp of the market price with only modest, similarly-reasoned discounts, showing insufficient differentiation given the acknowledged high uncertainty (no direct current Coding-subcategory read); a wider spread or more explicit sensitivity analysis on the "no clean confirmation" gap would better reflect true epistemic uncertainty rather than both anchoring near-identically to Kalshi.
6. Aggregator Opus · final calibrated forecast
85% (59% confidence)
Yes: 85% No: 15%
The Kalshi anchor (90% YES, +40pp/30d) plus a very short remaining horizon (~2 weeks) makes Anthropic the clear favorite: Claude Opus 5's late-July coding gains landed, Anthropic leads LiveBench overall and BenchLM's composite coding ranking, and the most credible challenger (Gemini 3.5 Pro) is reportedly months behind schedule with Google only shipping mid-tier Flash models in August. However, the critique lands on real weaknesses: no source directly confirms the LiveBench *Coding* sub-category ranking, OpenAI held that exact category at 99% as recently as Jan-2026, the BenchLM gap is razor-thin (81.1 vs 78.7), and LiveBench's monthly contamination-control question rotation injects exogenous variance right before the Aug-31 snapshot. The naive monthly-flip base rate is not the right model over a 2-week window — leadership is empirically sticky between frontier releases — so it justifies only a modest, not dramatic, discount. I therefore shade a few points below both the market and Forecast 1, landing at 0.85, which also implicitly leaves room for tail flips from OpenAI's pending GPT-5.x update or an unexpected xAI/open-weight benchmark spike.
Pipeline Timing
Total pipeline time: 238.1s
Per-tool research timings shown in the Research section above.