# Current state
Meta has abandoned the Llama brand for its frontier models, releasing "Muse Spark" (April 2026) and "Muse Spark 1.1" (July 2026) via Meta Superintelligence Labs. Meta self-reports a "Contemplating" multi-agent/tool-use mode scoring 58% on HLE — above the 55% threshold — but independent trackers (Artificial Analysis, Scale AI official leaderboard) put Meta's verified/standard score at ~40-51%, below threshold and below the current independently-verified frontier SOTA (Claude Fable 5, 55.5%, as of Aug 2026).
# Timeline of key events
- 2025-04: Llama 4 released; HLE score ~5-8% (confirmed, layerlens.ai/lmmarketcap).
- 2025-12-09/11: Reports emerge Meta delaying "Avocado" (next-gen model) to Q1 2026 amid reorg, Chris Cox exit (reported, cnbc.com/digitimes.com).
- 2026-04-08: Meta launches Muse Spark (codename "Avocado"), first Meta Superintelligence Labs model, replacing Llama line (confirmed, buildfastwithai.com, Wikipedia).
- 2026-06-04/05: Reports Meta "repeatedly pushes back" a new AI model release for developers (reported, WSJ via thehindu/dunyanews).
- 2026-07-08/09: Meta releases Muse Spark 1.1, unveiled by Zuckerberg; Meta self-reports HLE 58% (with tools/Contemplating mode) vs. independent Artificial Analysis measurement of 39.9% (confirmed release; self-reported score disputed).
- 2026-07-09: Commentators (Digg) flag conflict-of-interest concerns over Meta's self-reported Muse Spark 1.1 HLE benchmark claims (reported/opinion).
- 2026-08-11/16: Muse Glimmer (open model) and broader Meta AI product rollout continues; no new HLE record claimed (reported).
- 2026-08-22: Third-party leaderboard (Artificial Analysis) shows Claude Fable 5 leading HLE at 55.5%, ahead of all Meta figures on standardized measurement (confirmed per artificialanalysis.ai/pricepertoken.com).
# Event
Will any Meta-published model reach ≥55% HLE accuracy (per agi.safe.ai leaderboard) by Dec 31, 2026?
# Outcomes to forecast
- Yes (≥55%)
- No (<55%)
# Kalshi market anchor
Kalshi-direct price was NOT returned in this research pass (tool output missing). The closest available cross-market proxy is Polymarket, pricing this same underlying question at **49%** (7d: +2.5%, 30d: +10%, range 28-67.5% over 33 days, $15.3K volume) — trending upward but still near a coin flip. Treat Kalshi price as unknown/gap; Polymarket 49% is the best current consensus proxy.
# Sub-question answers
1. **Polymarket price for this/sibling thresholds?** Only this exact 55% threshold market was found (49%); no sibling 25%/40%/70% Meta-HLE markets identified on Polymarket [polymarket_direct/related].
2. **Highest Meta HLE score on agi.safe.ai to date?** Not directly confirmed on agi.safe.ai itself; via comparable leaderboards, Meta's Muse Spark self-reports 42.8% (no tools)/50.4% (with tools) and a disputed "58%" Contemplating-mode figure; independent Artificial Analysis measures only 39.9%; Scale AI's official leaderboard lists Meta's "Muse Spark" at 40.56% [claude_news, venturebeat, Wikipedia].
3. **Current cross-lab SOTA and its rise?** As of Aug 2026, Artificial Analysis leaderboard leader is Claude Fable 5 at 55.5%, followed by Claude Opus 5 variants (54.9%, 54.4%); Scale AI's stricter official leaderboard shows lower figures (Gemini 3.1 Pro 46.44%, GPT-5.4 Pro 44.32%). SOTA has risen from ~3% (Jan 2025) to ~41.5% (late 2025) to ~50-55% (mid-2026) [claude_news, code_execution extrapolation].
4. **Has Meta released a 2026 frontier model with HLE results?** Yes — Muse Spark (April 2026) and Muse Spark 1.1 (July 2026), both from Meta Superintelligence Labs, with published (self-reported) HLE scores [claude_news, gdelt_news, Wikipedia].
5. **Does Meta publish HLE evals / is it on agi.safe.ai?** Meta publishes its own benchmark claims (self-reported), but independent leaderboards (Artificial Analysis, Scale AI) separately measure and generally show materially lower results than Meta's self-reports; direct confirmation of agi.safe.ai inclusion not found in research, but Meta models appear on comparable/derivative leaderboards [claude_news].
6. **Shift to closed-source and timeline?** Confirmed — Meta abandoned Llama's open-weight strategy for Muse Spark (proprietary), amid leadership turmoil (Chris Cox exit) and repeated release delays reported through 2025-2026 [cnbc, digitimes, thehindu].
# Key facts (high-confidence, factual)
1. [Wikipedia] Meta Superintelligence Labs released Muse Spark in April 2026, replacing Llama.
2. [venturebeat/deeplearning.ai] Meta self-reports 58% HLE (Contemplating/tool-use mode); independent Artificial Analysis measured only 39.9%.
3. [artificialanalysis.ai via claude_news] Current independently-verified HLE leader (Aug 2026) is Claude Fable 5 at 55.5%, not Meta.
4. [Scale AI/Wikipedia] Official Scale AI leaderboard places Meta's Muse Spark at 40.56%, below GPT-5.4 Pro and Gemini 3.1 Pro.
5. [digitimes/cnbc] Meta delayed its flagship model release and underwent leadership reorg (Cox exit) in late 2025.
# Cross-market signals
- Kalshi related: No direct sibling HLE markets found; only tangential Meta business markets (headcount, DAP) unrelated to benchmark performance.
- Polymarket: 49% YES on this identical question, up from 28% low and trending toward 67.5% high over the past month — indicating rising but contested belief.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Commentators (Digg, July 2026) allege conflict-of-interest concerns over Meta's self-reported Muse Spark 1.1 HLE claims, suggesting the 58% figure is not credible/standardized.
- Code-execution model estimates ~15-20% probability of YES, citing Meta's persistent historical lag versus frontier labs as the binding constraint, despite ~67% likelihood frontier SOTA overall exceeds 55% by Dec 2026.
- Multiple outlets frame Muse Spark as competitive but not frontier-leading (trailing on coding/agentic benchmarks too).
# Directional lean per outcome
- **Yes**: Meta's own claimed 58% figure exists; Polymarket pricing near 49% shows real market uncertainty, not clear NO consensus; Meta is actively iterating (Spark→1.1) and could improve before Dec 2026.
- **No**: All independent/standardized measurements (Artificial Analysis, Scale AI) show Meta below 55% (39.9%-50.4%); resolution source (agi.safe.ai) likely uses standardized methodology, not self-reported multi-agent scores; Meta has a consistent history of lagging frontier by wide margins; conflict-of-interest concerns undermine Meta's self-reported figures.
# Gaps / unknowns
- No direct Kalshi YES price was retrieved for this ticker — a critical anchor gap.
- Unclear whether agi.safe.ai leaderboard would recognize Meta's "Contemplating" tool-augmented mode as a valid comparable HLE Accuracy entry, which is decisive for resolution.
- No confirmation of Meta's actual listed score directly from agi.safe.ai (research relied on secondary trackers with conflicting numbers).
- Unknown whether Meta will release a further/larger model in Q4 2026 before close.
# Calibration anchors
- Polymarket proxy price (best available): 49% YES, uptrending (+10% in 30 days).
- Independent benchmark data suggests true current Meta score sits ~15pp below threshold on standardized metrics; historical Meta HLE performance has consistently lagged SOTA by large margins (Llama 4: 5-8% vs. contemporaneous SOTA ~20-25%).
- Precedent: benchmark self-reports from labs (all vendors) often diverge substantially from third-party verification, typically favoring the vendor.