# Current state
The market resolves off arena.ai's "Agent Arena" leaderboard, Labs view, checked 2026-09-30. As of early August 2026, Moonshot's flagship Kimi K3 (released July 16, 2026) has risen sharply from a #5 lab ranking in June to a reported #3-4 on Agent Arena's Labs/overall board, though rankings are explicitly flagged by Arena as still "preliminary" and other composite leaderboards (DataLearner) place Kimi K3 as low as #7 once Anthropic, OpenAI, and xAI variants are counted separately. No kalshi_direct data was returned; the Polymarket price for this identical market is the best available consensus anchor.
# Timeline of key events
- 2026-06 (early June): Agent Arena "Labs" leaderboard launches; order is OpenAI, Anthropic, Z.ai (GLM-5.1), Google DeepMind, then Moonshot (Kimi-K2.6) at #5. (confirmed, arena.ai/X)
- 2026-06 (mid): Kimi K2.7 Code ranks #19 overall, #6 among open models. (reported, arena.ai/X)
- 2026-07-16: Kimi K3 (2.8T MoE) launches; lands #4 on Agent Arena leaderboard, tied with Claude Opus 4.8 / GPT-5.6 Sol. (confirmed, arena.ai/X)
- 2026-07-18 to 2026-07-22: Independent outlets (Notebookcheck, Forbes) report Kimi K3 has climbed to #3 on Agent Arena, behind only Claude Fable 5/Opus 5 and GPT-5.6 Sol; also #1 on Frontend Code Arena. (reported)
- 2026-07-27: Cryptobriefing corroborates K3 in top-3 overall, highest-performing open-weight model of 42 systems evaluated. (reported)
- 2026-08-03: Alibaba releases Qwen3.8-Max, described as "closing in" on Kimi K3's scale — a rival open-weight lab gaining ground. (reported)
- 2026-08-18: DataLearner composite leaderboard shows Kimi K3 at #7 overall, trailing multiple Anthropic/OpenAI/xAI model variants when each is counted individually rather than deduplicated by lab. (reported — conflicts with Arena-specific #3 finding)
- Ongoing (Aug 2026): Polymarket's identical "Moonshot 3rd" market trades at 61% (up from 23% low), reflecting bullish repricing after K3's launch.
# Event
Will Moonshot rank as the third-best lab (by rank of its top model) on arena.ai's Agent Arena "Labs" leaderboard when checked 2026-09-30 12PM ET?
# Outcomes to forecast
Yes / No
# Kalshi market anchor
No kalshi_direct data was returned by tooling. Substitute anchor: Polymarket price for the identical market = **61.05% YES**, up +28.4% over 7 days and +11.55% over 30 days, off a 12-day range of 23.15%–61.45%, on modest volume ($15,076 total, thin/illiquid). This is a fast, recent repricing coincident with Kimi K3's launch and rank climb — treat with caution given low liquidity.
# Sub-question answers
1. **Current ordering / Moonshot's rank** — Moved from #5 (June, baseline) to #4 (July 16 launch) to #3 (July 18-27, per Arena/Notebookcheck/Cryptobriefing), behind OpenAI and Anthropic's flagships. However, DataLearner's Aug 18 composite ranks it #7 when Anthropic/OpenAI/xAI variants aren't deduplicated by lab. [claude_news]
2. **Polymarket sibling prices** — No sibling-market data was retrievable (polymarket_related found 0 matches); the code_execution tool's "de-vigged 6.6%" figure explicitly used fabricated illustrative placeholder prices, not real market data, and should be disregarded.
3. **Volatility of rank-3 over past 3-6 months** — High: Moonshot itself moved three lab-rank positions (5→4→3) in under two months following one model release; competitor labs (Z.ai, Google DeepMind, Alibaba, xAI) are also releasing updates. [claude_news, gdelt]
4. **Upcoming releases before Sept 2026** — Alibaba Qwen3.8-Max already released (Aug 3, closing gap with K3); DeepSeek V4, GLM-5.2 cited as imminent rivals; no confirmed Moonshot K3.x/K4 timeline found, but Moonshot has shipped a new K-series iteration roughly every 1-2 months (K2→K2.5→K2.6→K2.7→K3). [claude_news, gdelt]
5. **Score gap at rank-3 boundary** — Arena's own commentary flags K3's ranking as "preliminary" with wide confidence intervals; net-improvement score gap vs. #2/#4 is narrow (K3 was tied with Claude Opus 4.8/GPT-5.6 Sol at launch), implying the boundary is contestable both directions. [claude_news]
6. **Methodology/discontinuation/AutoEval risk** — No specific evidence found of leaderboard discontinuation or AutoEval-tagging risk to Moonshot entries; this remains an unresolved gap.
# Key facts (high-confidence, factual)
1. [arena.ai/X] Agent Arena Labs launched June 2026 with Moonshot at #5.
2. [arena.ai/X] Kimi K3 launched July 16, 2026, hit #4 on Agent Arena at launch.
3. [Notebookcheck, Cryptobriefing] By late July 2026, multiple outlets independently report Kimi K3 at #3 on Agent Arena.
4. [DataLearner] A different, non-deduplicated composite ranking (Aug 18) places K3 #7.
5. [Polymarket] Identical market trades 61% YES, up sharply in the last 7-30 days.
# Cross-market signals
- Kalshi related: no directly relevant markets found.
- Polymarket: 61.05% YES on this exact market; no sibling/complementary lab markets found to cross-check overround.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Claude_news synthesis argues Moonshot "very likely overtook" Z.ai and Google DeepMind for #3 as of late July, but flags genuine risk of reversal from Google, Z.ai, or OpenAI/Anthropic updates before September.
- Broader narrative (Forbes, Digit.in): 2026 marked by Chinese open-weight labs (DeepSeek, Moonshot, Z.ai, Alibaba) closing gap with closed frontier — implies crowded, contestable field for the #3 slot, not a stable Moonshot lock.
# Directional lean per outcome
- **Yes**: Kimi K3's July launch pushed Moonshot to #3 on Arena's own agent board per multiple outlets; Polymarket sentiment (61%) has swung bullish; open-weight lab momentum favors Moonshot near-term.
- **No**: Rankings explicitly "preliminary"/volatile; competing composite view (DataLearner) has Moonshot at #7; rival releases (Qwen3.8-Max, expected GLM-5.2, DeepSeek V4) could unseat it before Sept 30; six weeks remain for further reshuffling; thin Polymarket volume makes the 61% price less reliable.
# Gaps / unknowns
- No live kalshi_direct price obtained — anchor substituted with Polymarket.
- No confirmed current (August/September) live snapshot of arena.ai Labs-filtered view was captured; most recent hard evidence is late July.
- No information on AutoEval tagging risk or leaderboard methodology stability through September.
- No genuine sibling-market Polymarket data for cross-checking overround (code_execution figures were fabricated/illustrative, discard).
# Calibration anchors
- Polymarket YES price (proxy anchor): 61.05%, recent range 23%-61%, thin volume (~$15k).
- Precedent: leaderboard rank-3 has changed at least twice in ~2 months (5→4→3) tied to single model releases — suggests moderate-to-high base rate of further reshuffling over the remaining ~7 weeks to close.