# Current state
Resolution hinges on the official agi.safe.ai HLE leaderboard listing a Moonshot Kimi model at ≥50% "HLE Accuracy" by Dec 31, 2026. No research source directly confirms whether Kimi models currently appear on agi.safe.ai itself; all figures found come from Moonshot's own blog and third-party trackers (Artificial Analysis, llm-stats, benchlm), which is a key ambiguity. Self-reported Kimi HLE numbers have historically run ~2x higher than independently verified figures, and a later 50.2% figure for K2.5 traces to a single low-reliability secondary source (codecademy.com), not the resolution source.
# Timeline of key events
- 2025-01: HLE benchmark launches; frontier scores near ~9% (code_execution analysis).
- 2025-11: Kimi K2 Thinking released; Moonshot self-reports 44.9% HLE "with tools" (search/python/web-browsing), claimed SOTA at the time (confirmed — kimi.ai blog).
- 2025-11-09: Independent trackers report much lower Kimi scores: Artificial Analysis measures 22.3% (no tools); Zvi Mowshowitz cites 23.9% (confirmed — artificialanalysis.ai, thezvi.substack.com).
- 2026-01-27: Kimi K2.5 released (1T-param, "Agent Swarm" multimodal model); one secondary source claims 50.2% HLE (reported, single low-confidence source — codecademy.com).
- ~2026-04: Kimi K2.6 released (reported — mysummit.school).
- 2026-mid: Kimi K2.7-Code released (reported — mysummit.school).
- 2026-07-16: Kimi K3 released (2.8T-param MoE, largest open-weights model; ranked 2nd of 47 models in independent testing behind GPT-5.6 Sol) (confirmed release, benchmark ranking reported — Wikipedia/mysummit.school).
- Ongoing (as of research date): Third-party leaderboards (Artificial Analysis, llm-stats, benchlm) show top HLE scores 55–65%, led by Anthropic (Claude Fable 5: 55.5%, Opus 5: 54.9%) and OpenAI; none list a verified Kimi entry ≥50% (reported).
# Event
Will any Moonshot Kimi model reach ≥50% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026?
# Outcomes to forecast
- Yes
- No
# Kalshi market anchor
No direct kalshi_direct data returned; the only same-ticker market data available is from Polymarket: current price 45% ("Yes"), down sharply from a 7-day high (7d change -11.5%, 30d change -3.0%), range 24.5%–60.5% over 41 days, total volume ~$15,161. This is the best available consensus anchor and shows recent bearish momentum after an earlier peak near 60%.
# Sub-question answers
1. **Highest Kimi HLE score to date / on official leaderboard?** Self-reported: K2 Thinking 44.9% (with tools, Nov 2025); independently measured only 22.3–23.9%. A later, unverified claim puts K2.5 at 50.2% (Jan 2026, single low-reliability source). No confirmation these appear on agi.safe.ai itself — [kimi.ai, artificialanalysis.ai, thezvi.substack.com, codecademy.com].
2. **Cross-lab SOTA and pace?** Current SOTA per Artificial Analysis is Claude Fable 5 (55.5%), Opus 5 (54.9%); frontier tool-augmented scores rose from ~9% (Jan 2025) to ~44.5% (Nov 2025), ~3.55pp/month, with likely deceleration afterward [artificialanalysis.ai, code_execution].
3. **Tool augmentation?** Yes — Moonshot's reported 44.9% explicitly used search/python/web tools. Unclear whether the official leaderboard labels such agentic scores as canonical "HLE Accuracy" — unresolved ambiguity [kimi.ai].
4. **Is agi.safe.ai actively listing Kimi?** Not directly confirmed by research; only third-party trackers were checked, none showing a Kimi entry ≥50%. Risk exists that Kimi scores may not be promptly/officially listed — a structural gap in evidence.
5. **Future Kimi releases in 2026?** Confirmed cadence: K2.5 (Jan), K2.6 (~Apr), K2.7-Code, K3 (Jul, 2.8T params, 2nd of 47 in independent testing) — Moonshot maintains ~quarterly release cadence with large jumps [Wikipedia, mysummit.school].
6. **Gap to 50% and required pace?** From independently-verified baseline (~22–24%), gap is ~26–28pp, requiring faster-than-historical no-tool growth (~1.6pp/month, insufficient). From self-reported tool-augmented baseline (~44.9%), gap is only ~5pp, achievable within one release cycle given historical 10–15pp jumps between Kimi generations [code_execution].
# Key facts (high-confidence, factual)
1. [kimi.ai] K2 Thinking self-reported 44.9% HLE with tools (Nov 2025).
2. [artificialanalysis.ai] Independent verification of K2 Thinking: 22.3% (no tools).
3. [Wikipedia] Kimi K3 (July 2026) is a 2.8T-param model, largest open-weights model to date.
4. [artificialanalysis.ai] Current cross-lab SOTA (non-Kimi) ~55.5% (Claude Fable 5).
5. [mysummit.school] K3 ranked 2nd of 47 models in independent testing (source reliability moderate/unverified).
# Cross-market signals
- Polymarket (same ticker): 45% Yes, down from 60.5% high, trending down.
- No Kalshi-specific or Polymarket-related HLE/Kimi markets found (only an unrelated esports "HLE" match).
- No sportsbook signal available (not applicable).
# Analyst opinions and speculation
- Zvi Mowshowitz and Artificial Analysis flag a pattern of Moonshot overstating HLE results by ~2x versus independent verification.
- code_execution model: point estimate ~55–65% probability of Yes if tool-augmented scores count as the resolving metric; drops to ~20–30% if strict text-only scoring is enforced.
# Directional lean per outcome
- **Yes**: Self-reported trajectory (44.9%→claimed 50.2%) already near/over threshold; rapid Moonshot release cadence (K2.5/K2.6/K2.7/K3) and industry-wide upward trend support crossing 50% under tool-augmented framing.
- **No**: Independent/official verification consistently ~2x lower than self-reports; no third-party leaderboard (Artificial Analysis, llm-stats, benchlm) currently shows any Kimi model ≥50%; unclear if agi.safe.ai lists Kimi at all; Polymarket pricing has been falling toward 45%, suggesting market skepticism.
# Gaps / unknowns
- No direct check of agi.safe.ai leaderboard content for Kimi models.
- Reliability of the 50.2% K2.5 claim (single tertiary source) is low.
- Whether official leaderboard would classify tool-augmented scores as "HLE Accuracy" per market rules is unresolved.
# Calibration anchors
- Polymarket same-ticker price: 45% Yes (recent high 60.5%, declining).
- Historical precedent: self-reported AI benchmark claims from Chinese labs have often been ~1.5–2x higher than independently verified figures (K2 Thinking case).