# Current state
The market resolves on the official agi.safe.ai/Scale AI HLE leaderboard, where the current best documented Meta entry (Muse Spark, MSL's post-Llama flagship) sits at ~40-45% per independent trackers — well below the 55% bar. Meta's own marketing claims of 50.2%-58% rely on an ensemble "Contemplating mode" not obviously equivalent to the standard "HLE Accuracy" leaderboard metric, and are contested by independent evaluators (Artificial Analysis) and press (conflict-of-interest allegations).
# Timeline of key events
- 2025-04: Meta releases Llama 4 (Scout/Maverick); best documented HLE score ~6.2% (rank ~98/97) — confirmed low benchmark, per LayerLens/BenchmarkList (claude_news).
- Late 2025–early 2026: Llama 4 Behemoth (2T param) finishes training but is shelved due to "poor internal performance"; never formally released or cancelled (reported, NYT via claude_news).
- 2026-04-08: Meta Superintelligence Labs (led by Alexandr Wang) launches Muse Spark, a closed-weight model replacing the Llama line (confirmed, Wikipedia + officechai.com + claude_news).
- 2026-04: Muse Spark 1.0 launch — Meta self-reports 50.2% HLE (no tools, "Contemplating" mode) and a separate 58% figure for full Contemplating mode; independent Artificial Analysis measures Thinking-mode score at 39.9% (reported/disputed).
- 2026-07-09: Commentators (Digg) allege conflict-of-interest issues compromising Meta's Muse Spark 1.1 HLE benchmark claims (reported).
- 2026-07 (undated, ~mid): Muse Spark 1.1 released; independent HLE ≈45% (Artificial Analysis), Intelligence Index 51 (reported).
- 2026-08-11: Meta unveils Muse Glimmer, a smaller 30B local-run model (not a frontier/HLE contender) (confirmed, Forbes).
- 2026-08-16: Official Scale AI public leaderboard snapshot: Gemini 3.1 Pro (thinking high) 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Gemini 3 Pro 37.52% (confirmed per IntuitionLabs summary of agi.safe.ai partner site).
- 2026-08 (~"Muse Spark 1.2"): Independent HLE score dips to ~44% (from 45%), Intelligence Index rises to 54 (reported, Artificial Analysis).
- 2026-09-02: Frontier field leaders (non-Meta) hit 55-65% on various trackers — Claude Fable 5.1 at 65% (BenchLM) / 59.1% (Artificial Analysis) — confirming the 55%+ bar is achievable technically but Meta remains ~15-20 points behind (confirmed).
# Event
Will any Meta-published model reach ≥55% accuracy on the official Humanity's Last Exam (agi.safe.ai) leaderboard by Dec 31, 2026?
# Outcomes to forecast
- Yes (Meta model hits ≥55% HLE)
- No (Meta stays below 55%)
# Kalshi market anchor
Current YES price: **44%** (Polymarket data, same ticker). Trend: -5.5% over 7 days, -2% over 30 days — declining. Range has been wide (28%–67.5%) over 42 days of trading, volume modest (~$22.9K total). Market has cooled from earlier optimism, consistent with underwhelming independent Muse Spark verification scores.
# Sub-question answers
1. **Current top HLE score / 12-month trend** — Independent trackers show frontier leaders now at 55-65% (Claude Fable 5.1: 65% per BenchLM, 59.1% per Artificial Analysis, Sept 2026), up from single digits in early 2025 (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — an unprecedented ~1-year climb led by Anthropic/OpenAI/Google. The official Scale AI leaderboard is lower/lagged, topping out near 46% (Gemini 3.1 Pro, Aug 2026).
2. **Best Meta score / leaderboard presence** — Meta's Muse Spark appears on the official Scale AI leaderboard at 40.56% (Aug 2026, IntuitionLabs). Independent AA trackers put Muse Spark 1.1/1.2 at 44-45%. Meta's own marketing claims 50.2-58% via a multi-agent "Contemplating mode," disputed by independent evaluators and press (conflict-of-interest allegations, Digg).
3. **2026 frontier model plans** — Llama 4 Behemoth is effectively dead/shelved (never shipped). Meta pivoted fully to closed-weight Muse Spark line under MSL (launched April 2026), with iterative updates (1.0→1.1→1.2) showing incremental, sometimes regressing, gains (Intelligence Index 43→51→54). Muse Glimmer (Aug 2026) is a smaller local model, not a benchmark contender.
4. **Tool-augmented/agentic configs** — Meta's highest self-reported figures (58%) come from an ensemble "Contemplating mode," which functions similarly to tool/agentic augmentation; independent verifiers don't replicate this in single-model "Thinking mode" tests (39.9-45%). Whether the official leaderboard would accept such a mode as "HLE Accuracy" is unclear/disputed — a key resolution risk.
5. **Gap vs. OpenAI/Google/Anthropic/xAI** — Meta trails significantly: on the official leaderboard, Meta (40.56%) sits below GPT-5.4 Pro (44.32%) and Gemini 3.1 Pro (46.44%), and far below Anthropic's ~60-65% frontier scores. Meta also lags on coding/hard reasoning generally (Llama 4 "trails DeepSeek V4 and Qwen 3.5/3.6"). No base rate exists for Meta closing a 15-20pt benchmark gap in under 4 months given repeated underperformance vs. self-reported claims.
6. **Related markets pricing** — No other HLE-threshold or lab-specific Kalshi/Polymarket markets found (0 matches on Polymarket related-market scan); this ticker is the only relevant HLE market identified.
# Key facts (high-confidence, factual)
1. [Polymarket] Current YES price 44%, declining trend.
2. [IntuitionLabs/agi.safe.ai] Official leaderboard: Muse Spark at 40.56% (Aug 16, 2026); top model (Gemini 3.1 Pro) at 46.44%.
3. [Wikipedia] Muse Spark replaced Llama as Meta's flagship in April 2026 under MSL.
4. [claude_news/NYT] Behemoth shelved, never released.
5. [Artificial Analysis via claude_news] Independent Muse Spark scores: 1.0 39.9%, 1.1 ~45%, 1.2 ~44% — plateauing below 55%, not converging upward.
# Cross-market signals
- Kalshi related: No direct HLE markets found besides this ticker; tangential Meta markets (headcount, antitrust) unrelated to AI benchmarks.
- Polymarket: Same ticker mirrored — 44% YES, high volatility (28-67.5% range), suggesting past optimism (possibly around Meta's 58% self-reported claim) has since faded as independent verification came in lower.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Claude_news synthesis: reaching 55%+ "would require an enormous, unprecedented leap...inconsistent with Meta's current trajectory."
- Digg commentary alleges conflict-of-interest in Meta's self-reported benchmark methodology, casting doubt on Meta's higher claimed figures being accepted for official resolution.
# Directional lean per outcome
- **Yes**: Field-wide progress is rapid (frontier leaders already exceed 55-65%); Meta's self-reported "Contemplating mode" claims 58%, and if such ensemble scoring were credited on the official leaderboard, Yes could resolve. Four more months of releases remain.
- **No** (favored): Official leaderboard has Meta at ~40.56%; independent trackers show a plateau (~44-45%) across three iterative releases with no clear upward trajectory; flagship Behemoth is dead; Meta trails all four major competitors by double digits.
# Gaps / unknowns
- Whether "Contemplating mode" ensemble scores would count as official "HLE Accuracy" per resolution rules is unresolved and pivotal.
- No confirmation of further Meta frontier releases planned before Dec 31, 2026.
- Official agi.safe.ai leaderboard update cadence/lag vs. third-party trackers unclear.
# Calibration anchors
- Kalshi/Polymarket current YES price: 44% (anchor), down from a 67.5% high.
- Precedent: Meta's Muse Spark line has repeatedly underperformed its own marketing claims by 10-18 points when independently verified (1.0: 58% claimed vs 39.9% verified).