← Back to scans

Will the highest score achieved by a Moonshot Kimi model on Humanity’s Last Exam in 2026 be 50% or higher?

0x0c49d3e9e2ca6bcf8a876a640948abe8a23823a9379ca942d0612704c22ccdb4 · Science and Technology · 2026-09-01
37%
Agent
45%
Market Price
-8.0%
Edge
50%
Confidence
Volume: 15,161
Spread: 4.0c
Days to resolution: 121
Markets in event: 5
Final Rationale
Both forecasts landed at 0.42, but the critique correctly identifies that multiple independent bearish factors compound: no confirmation Kimi appears on the resolution source (agi.safe.ai) at all, unresolved whether tool-augmented scores count as canonical 'HLE Accuracy', and Moonshot's documented ~2x self-report inflation (verified 22-24% vs claimed 44.9%) which implies the single-source 50.2% K2.5 claim likely reflects a true verified score far below threshold. Critically, as of mid-to-late 2026 no checked third-party leaderboard shows any Kimi entry ≥50% despite K2.5, K2.6, K2.7 and K3 having already shipped — this is late-cycle evidence, not just early-year uncertainty. The K3 tail (2nd of 47 in independent testing) and remaining ~5 months keep Yes meaningfully alive, and the declining Polymarket price (60%→45%) suggests the market itself is absorbing these same negatives. I move modestly below both forecasts and the anchor to 0.37 Yes.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 1$ follow-up
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct claude_news gdelt_news polymarket_related kalshi_related wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the highest HLE accuracy any Moonshot Kimi model (e.g., Kimi K2, K2 Thinking) has achieved to date, and is it listed on the official agi.safe.ai leaderboard?
  2. What is the current state-of-the-art HLE score across all labs, and how fast have top scores been improving per quarter since the benchmark launched in early 2025?
  3. Do Kimi's reported HLE results use tool/agentic augmentation, and does the official leaderboard list those augmented scores as 'HLE Accuracy'?
  4. Is the agi.safe.ai leaderboard actively maintained and does it include Moonshot/Kimi models, or would a strong Kimi score risk never appearing on the resolution source?
  5. Has Moonshot announced or been rumored to release a next-generation Kimi model (K3 or similar) in 2026, and what benchmark gains are expected?
  6. How large is the gap between Kimi's best current HLE score and 50%, and what improvement rate would be needed to close it by December 2026?
Planner reasoning
This is a Polymarket question about whether any Moonshot Kimi model reaches ≥50% HLE accuracy on the official leaderboard by end of 2026. The key drivers are the current best Kimi score, the overall frontier trajectory on HLE (top models globally), whether the agi.safe.ai leaderboard is actively updated with Kimi entries, and whether tool-augmented scores count. I'll anchor on the Polymarket price, search news for latest Kimi/HLE benchmark results, and triangulate with related markets on other HLE thresholds and rival labs.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a Moonshot Kimi model on Humanity’s Last Exam in 2026 be 50% or higher?** - Current price (probability): 45.00% - 7-day price change: -11.50% - 30-day price change: -3.00% - Total volume: $15,161 (USD notional) - Price range: 24.
claude_news OK 26.9s 6 **Key findings on Moonshot Kimi HLE scores:** - **Kimi K2 Thinking (Nov 2025)** achieved a self-reported HLE score of **44.9% with tools** (search, python, web-browsing), which Moonshot called "state-of-the-art" at the time, beating GPT-5's 41.7%. ( On Humanity's Last Exam K2 Thinking achieved a st
gdelt_news OK 100.2s 10 GDELT: 10 articles across 3 queries (lookback=60d). "Kimi K2 Humanity's Last Exam benchmark": error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28kimi+OR+humanity+OR+last+OR+exam+OR+benchmark%29+sourcelang%3Aenglish&mode=ArtList&f
polymarket_related OK 2.3s 1 Scanned 100 active Polymarket markets, kept 1 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Kimi': 0 markets | keyword 'Moonshot': 0 markets | keyword 'HLE': 1 markets
kalshi_related OK 2.1s 2 2 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Kimi': no matches
wikipedia OK 0.2s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 46.6s 0 **Key quantitative findings:** - **Historical frontier pace (tool-augmented SOTA):** HLE went from ~9% (Jan 2025) to ~44.5% (Nov 2025) → a linear rate of **~3.55 pp/month**. Naive linear extrapolation to Dec 2026 (+13 months) gives an unrealistic ~91%, so a decelerating (logistic-decay, half-life =
3. Evidence Brief Sonnet · 6361 chars
# Current state Resolution hinges on the official agi.safe.ai HLE leaderboard listing a Moonshot Kimi model at ≥50% "HLE Accuracy" by Dec 31, 2026. No research source directly confirms whether Kimi models currently appear on agi.safe.ai itself; all figures found come from Moonshot's own blog and third-party trackers (Artificial Analysis, llm-stats, benchlm), which is a key ambiguity. Self-reported Kimi HLE numbers have historically run ~2x higher than independently verified figures, and a later 50.2% figure for K2.5 traces to a single low-reliability secondary source (codecademy.com), not the resolution source. # Timeline of key events - 2025-01: HLE benchmark launches; frontier scores near ~9% (code_execution analysis). - 2025-11: Kimi K2 Thinking released; Moonshot self-reports 44.9% HLE "with tools" (search/python/web-browsing), claimed SOTA at the time (confirmed — kimi.ai blog). - 2025-11-09: Independent trackers report much lower Kimi scores: Artificial Analysis measures 22.3% (no tools); Zvi Mowshowitz cites 23.9% (confirmed — artificialanalysis.ai, thezvi.substack.com). - 2026-01-27: Kimi K2.5 released (1T-param, "Agent Swarm" multimodal model); one secondary source claims 50.2% HLE (reported, single low-confidence source — codecademy.com). - ~2026-04: Kimi K2.6 released (reported — mysummit.school). - 2026-mid: Kimi K2.7-Code released (reported — mysummit.school). - 2026-07-16: Kimi K3 released (2.8T-param MoE, largest open-weights model; ranked 2nd of 47 models in independent testing behind GPT-5.6 Sol) (confirmed release, benchmark ranking reported — Wikipedia/mysummit.school). - Ongoing (as of research date): Third-party leaderboards (Artificial Analysis, llm-stats, benchlm) show top HLE scores 55–65%, led by Anthropic (Claude Fable 5: 55.5%, Opus 5: 54.9%) and OpenAI; none list a verified Kimi entry ≥50% (reported). # Event Will any Moonshot Kimi model reach ≥50% HLE Accuracy on the official agi.safe.ai leaderboard by Dec 31, 2026? # Outcomes to forecast - Yes - No # Kalshi market anchor No direct kalshi_direct data returned; the only same-ticker market data available is from Polymarket: current price 45% ("Yes"), down sharply from a 7-day high (7d change -11.5%, 30d change -3.0%), range 24.5%–60.5% over 41 days, total volume ~$15,161. This is the best available consensus anchor and shows recent bearish momentum after an earlier peak near 60%. # Sub-question answers 1. **Highest Kimi HLE score to date / on official leaderboard?** Self-reported: K2 Thinking 44.9% (with tools, Nov 2025); independently measured only 22.3–23.9%. A later, unverified claim puts K2.5 at 50.2% (Jan 2026, single low-reliability source). No confirmation these appear on agi.safe.ai itself — [kimi.ai, artificialanalysis.ai, thezvi.substack.com, codecademy.com]. 2. **Cross-lab SOTA and pace?** Current SOTA per Artificial Analysis is Claude Fable 5 (55.5%), Opus 5 (54.9%); frontier tool-augmented scores rose from ~9% (Jan 2025) to ~44.5% (Nov 2025), ~3.55pp/month, with likely deceleration afterward [artificialanalysis.ai, code_execution]. 3. **Tool augmentation?** Yes — Moonshot's reported 44.9% explicitly used search/python/web tools. Unclear whether the official leaderboard labels such agentic scores as canonical "HLE Accuracy" — unresolved ambiguity [kimi.ai]. 4. **Is agi.safe.ai actively listing Kimi?** Not directly confirmed by research; only third-party trackers were checked, none showing a Kimi entry ≥50%. Risk exists that Kimi scores may not be promptly/officially listed — a structural gap in evidence. 5. **Future Kimi releases in 2026?** Confirmed cadence: K2.5 (Jan), K2.6 (~Apr), K2.7-Code, K3 (Jul, 2.8T params, 2nd of 47 in independent testing) — Moonshot maintains ~quarterly release cadence with large jumps [Wikipedia, mysummit.school]. 6. **Gap to 50% and required pace?** From independently-verified baseline (~22–24%), gap is ~26–28pp, requiring faster-than-historical no-tool growth (~1.6pp/month, insufficient). From self-reported tool-augmented baseline (~44.9%), gap is only ~5pp, achievable within one release cycle given historical 10–15pp jumps between Kimi generations [code_execution]. # Key facts (high-confidence, factual) 1. [kimi.ai] K2 Thinking self-reported 44.9% HLE with tools (Nov 2025). 2. [artificialanalysis.ai] Independent verification of K2 Thinking: 22.3% (no tools). 3. [Wikipedia] Kimi K3 (July 2026) is a 2.8T-param model, largest open-weights model to date. 4. [artificialanalysis.ai] Current cross-lab SOTA (non-Kimi) ~55.5% (Claude Fable 5). 5. [mysummit.school] K3 ranked 2nd of 47 models in independent testing (source reliability moderate/unverified). # Cross-market signals - Polymarket (same ticker): 45% Yes, down from 60.5% high, trending down. - No Kalshi-specific or Polymarket-related HLE/Kimi markets found (only an unrelated esports "HLE" match). - No sportsbook signal available (not applicable). # Analyst opinions and speculation - Zvi Mowshowitz and Artificial Analysis flag a pattern of Moonshot overstating HLE results by ~2x versus independent verification. - code_execution model: point estimate ~55–65% probability of Yes if tool-augmented scores count as the resolving metric; drops to ~20–30% if strict text-only scoring is enforced. # Directional lean per outcome - **Yes**: Self-reported trajectory (44.9%→claimed 50.2%) already near/over threshold; rapid Moonshot release cadence (K2.5/K2.6/K2.7/K3) and industry-wide upward trend support crossing 50% under tool-augmented framing. - **No**: Independent/official verification consistently ~2x lower than self-reports; no third-party leaderboard (Artificial Analysis, llm-stats, benchlm) currently shows any Kimi model ≥50%; unclear if agi.safe.ai lists Kimi at all; Polymarket pricing has been falling toward 45%, suggesting market skepticism. # Gaps / unknowns - No direct check of agi.safe.ai leaderboard content for Kimi models. - Reliability of the 50.2% K2.5 claim (single tertiary source) is low. - Whether official leaderboard would classify tool-augmented scores as "HLE Accuracy" per market rules is unresolved. # Calibration anchors - Polymarket same-ticker price: 45% Yes (recent high 60.5%, declining). - Historical precedent: self-reported AI benchmark claims from Chinese labs have often been ~1.5–2x higher than independently verified figures (K2 Thinking case).
4. Forecasters Opus + GPT · in parallel
claude-fable-5
0.42
Yes 42%
No 58%
The market anchor (Polymarket 45% Yes, trending down from 60%) reflects genuine uncertainty. Independently verified Kimi scores (22-24% for K2 Thinking) run ~2x below Moonshot's self-reported 44.9%, and no third-party leaderboard currently lists any Kimi model at ≥50%; the 50.2% K2.5 claim comes from a single low-reliability source. Moonshot's rapid release cadence (K2.5 through K3, with K3 ranked 2nd of 47 in independent testing) and industry-wide progress toward 55%+ leave a real path to Yes, but resolution requires the official agi.safe.ai listing, which is unconfirmed to even include Kimi models and may not credit tool-augmented scores. I lean slightly below the anchor given the persistent verification gap and bearish market momentum.
gpt-5.6-sol
0.57
Yes 42%
No 58%
The 45% Polymarket Yes price is the best available consensus anchor, but the lack of any confirmed Kimi score on the official agi.safe.ai leaderboard warrants a modest downward adjustment. Kimi's 44.9% tool-assisted self-report, rapid 2026 release cadence, and K3's strong independent ranking make a five-point improvement plausible if agentic scoring counts. However, independently measured Kimi scores have been only 22.3–23.9%, the claimed 50.2% K2.5 result is weakly sourced, and no checked third-party leaderboard shows Kimi at or above 50%. The unresolved official-listing and tool-eligibility questions therefore leave No slightly favored.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters converge on identical 0.42/0.58 outputs despite ostensibly independent reasoning, and neither explicitly grapples with the compounding effect of multiple independent uncertainties (whether Kimi is even listed on agi.safe.ai AND whether tool-augmented scores count AND whether the 50.2% claim is real) — stacking these should push probability lower than a single "modest adjustment" from the anchor suggests. 2. Neither forecast adequately addresses the structural gap flagged in the brief: "No confirmation these appear on agi.safe.ai itself" — if Kimi isn't even listed on the resolution source, this is arguably a near-automatic No regardless of actual score, and this binary listing risk deserves more explicit weight than a marginal haircut. 3. Both anchor heavily on Polymarket's 45% while treating it as a stable consensus, but the brief notes it's "down sharply from a 7-day high" with "recent bearish momentum after an earlier peak near 60%" — this trend could reflect the market resolving exactly the ambiguity (tool-augmented vs. text-only, listing status) that both forecasters flag as unresolved, suggesting the trend itself is informative and should pull probability further down rather than just anchoring near the current point price. 4. The code_execution model's bimodal estimate (55-65% if tool-augmented counts vs. 20-30% if strict text-only) is mentioned in the brief but neither forecaster explicitly weights these two scenarios or states a probability for which resolution standard is more likely — this is a critical missing sub-calculation that could materially shift the final number in either direction. 5. Neither forecast discusses the low reliability of the K2.5 50.2% source (single codecademy.com citation) with appropriate skepticism-adjusted weight — both treat it as a soft positive signal for Yes, but given Moonshot's established 2x self-report inflation pattern, a claimed 50.2% likely implies a true/verified score closer to 25-30%, which should be a stronger bearish signal than either forecast reflects. 6. Neither considers the tail scenario where K3 (2.8T params, "2nd of 47 models") could independently push scores well past 50% by year-end given its strong ranking, nor the opposite tail where agi.safe.ai simply never lists any Kimi model at all (making the question resolve No by default) — both tails are underexplored despite being flagged in the sub-question answers.
6. Aggregator Opus · final calibrated forecast
37% (50% confidence)
Yes: 37% No: 63%
Both forecasts landed at 0.42, but the critique correctly identifies that multiple independent bearish factors compound: no confirmation Kimi appears on the resolution source (agi.safe.ai) at all, unresolved whether tool-augmented scores count as canonical 'HLE Accuracy', and Moonshot's documented ~2x self-report inflation (verified 22-24% vs claimed 44.9%) which implies the single-source 50.2% K2.5 claim likely reflects a true verified score far below threshold. Critically, as of mid-to-late 2026 no checked third-party leaderboard shows any Kimi entry ≥50% despite K2.5, K2.6, K2.7 and K3 having already shipped — this is late-cycle evidence, not just early-year uncertainty. The K3 tail (2nd of 47 in independent testing) and remaining ~5 months keep Yes meaningfully alive, and the declining Polymarket price (60%→45%) suggests the market itself is absorbing these same negatives. I move modestly below both forecasts and the anchor to 0.37 Yes.
7. Entry Decision Opus Trader
Side
NO
Position Size
$500
Contracts
877
Trader Confidence
50%
Trade Rationale
This is a direction-agreement edge (bot 37% vs market 45%, both below 50%) with only 8pp of edge, which is historically the weakest setup class — so it needs concrete, recent evidence to clear the bar. It has some: as of late in the model cycle, no Kimi entry appears on the resolution source (agi.safe.ai) at all despite K2.5 through K3 having shipped, Moonshot has a documented ~2x self-report inflation pattern (verified 22-24% vs claimed 44.9%), and the sole 50.2% claim traces to a single low-reliability citation. The Devil's Advocate actually pushes further toward NO, arguing the compounding of listing risk × resolution-standard ambiguity × score inflation deserves more than a marginal haircut. The main caution is the bot's own admission that the 60%→45% price decline shows the market is absorbing these same negatives, and the unresolved bimodal scenario (55-65% if tool-augmented scores count) leaves real YES tail risk — hence minimum size despite a defensible thesis.
Allocation Logic
$500 — the floor — because this is a small-magnitude direction-agreement edge (historically ~36% hit rate), forecaster confidence is middling at 0.49, and the market's recent downtrend suggests much of the bearish case is already being priced in; the concrete listing-risk and inflation evidence justifies entry but not conviction sizing.
Entry price: $0.57
Current: $0.53
Status: OPEN
P&L: -$35.09
Pipeline Timing
Total pipeline time: 193.8s
Per-tool research timings shown in the Research section above.