← Back to scans

Will the highest score achieved by a Meta model on Humanity’s Last Exam in 2026 be 55% or higher?

0xb26b135b38bcbe4c7abe33467a8da10030286d29c4f662ca269ed3ae05db9cbb · Companies · 2026-08-24
40%
Agent
49%
Market Price
-9.0%
Edge
44%
Confidence
Volume: 15,286
Spread: 4.0c
Days to resolution: 129
Markets in event: 4
Final Rationale
Every independently standardized measurement of Meta's best 2026 models sits well below the bar (Artificial Analysis 39.9%, Scale AI 40.56%, Meta's own no-tools figure 42.8%), and Meta has a consistent history of trailing the frontier, which is itself only just at ~55.5%. The main YES pathways are (a) agi.safe.ai crediting Meta's self-reported 58% 'Contemplating' tool-augmented entry — the official HLE leaderboard has historically included tool/search-augmented rows, which makes this a live, quasi-binary branch — and (b) a Q4 Meta release (Spark 2.0) clearing 55% on standardized eval given Meta's rapid April→July iteration cadence. Because these two pathways are more than a minor haircut, I nudge above both forecasters' ~35% toward the Polymarket proxy (49%, uptrending), while staying below it since that market is thin, volatile, and the analytically grounded estimate points lower. Net: ~40% YES, reflecting genuine resolution-criteria ambiguity rather than a settled NO.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 9$ follow-ups
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price for this market and its sibling thresholds (e.g., 25%, 40%, 55%, 70%) for Meta on HLE in 2026?
  2. What is the highest HLE accuracy ever recorded by any Meta model (e.g., Llama 4, any Meta reasoning model) on the official agi.safe.ai leaderboard as of now?
  3. What is the current state-of-the-art HLE accuracy across all labs (GPT-5.x, Gemini 3, Claude Opus 4.5, Grok 4/5) and how fast has SOTA risen over the past 12 months?
  4. Has Meta announced or released a frontier reasoning model in 2026 (Llama 5 / Avalon / Meta Superintelligence Labs model), and has it published HLE benchmark results?
  5. Does Meta publish evaluations on HLE at all, and does the agi.safe.ai leaderboard include Meta models or only a curated set of frontier models?
  6. Has Meta signaled a shift toward closed-source/frontier model releases, and what is the reported timeline for its next major model?
Planner reasoning
This is a Polymarket question about whether any Meta model scores ≥55% on Humanity's Last Exam by end-2026. The key drivers are (a) whether Meta ships a frontier reasoning model in 2026 at all, (b) whether Meta models are even listed on the agi.safe.ai leaderboard, and (c) the current state-of-the-art HLE score trajectory across all labs. Market price on Polymarket is the primary anchor, supplemented by news on Meta Superintelligence Labs' roadmap and HLE SOTA progression.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a Meta model on Humanity’s Last Exam in 2026 be 55% or higher?** - Current price (probability): 49.00% - 7-day price change: +2.50% - 30-day price change: +10.00% - Total volume: $15,286 (USD notional) - Price range: 28.00% - 67.
polymarket_related OK 3.4s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'Meta model': 0 markets | keyword 'HLE': 0 markets | keyword 'Llama': 0 markets | keyword 'best AI model': 0 markets
kalshi_related OK 3.2s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'AI benchmark': ok | keyword 'Meta AI model': ok
claude_news OK 23.4s 8 Here are the key findings on HLE leaderboard standings and Meta's benchmarked scores: **Overall HLE leaderboard (top scores, as of Aug 2026):** - Artificial Analysis leaderboard: "Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) currently has the highest Humanity's Last Exam scor
claude_news OK 30.4s 14 Here are key findings on Meta's 2026 frontier model roadmap and HLE performance: - **Meta abandoned "Llama" branding for its frontier model, launching "Muse Spark" (internally "Avocado") on April 8, 2026** as the first model from Meta Superintelligence Labs (MSL), marking a shift from open-weight t
gdelt_news OK 61.0s 30 GDELT: 30 articles across 3 queries (lookback=90d). "Meta Llama Humanity's Last Exam benchmark": 10 hits | 'Meta Superintelligence Labs new model release': 10 hits | "Humanity's Last Exam record score": 10 hits
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
code_execution OK 69.6s 0 ## Key Findings **SOTA HLE trend extrapolation (two model families):** - **Logistic fit** to data points (Jan-25: 3%, mid-25: 25%, late-25: 41.5%, early-26: 50%) yields a natural ceiling of only **L≈53%**, projecting SOTA ≈ **53.2% (±5.6 std)** by Dec 2026 → P(SOTA > 55%) ≈ **0.37**. This model lik
3. Evidence Brief Sonnet · 7504 chars
# Current state Meta has abandoned the Llama brand for its frontier models, releasing "Muse Spark" (April 2026) and "Muse Spark 1.1" (July 2026) via Meta Superintelligence Labs. Meta self-reports a "Contemplating" multi-agent/tool-use mode scoring 58% on HLE — above the 55% threshold — but independent trackers (Artificial Analysis, Scale AI official leaderboard) put Meta's verified/standard score at ~40-51%, below threshold and below the current independently-verified frontier SOTA (Claude Fable 5, 55.5%, as of Aug 2026). # Timeline of key events - 2025-04: Llama 4 released; HLE score ~5-8% (confirmed, layerlens.ai/lmmarketcap). - 2025-12-09/11: Reports emerge Meta delaying "Avocado" (next-gen model) to Q1 2026 amid reorg, Chris Cox exit (reported, cnbc.com/digitimes.com). - 2026-04-08: Meta launches Muse Spark (codename "Avocado"), first Meta Superintelligence Labs model, replacing Llama line (confirmed, buildfastwithai.com, Wikipedia). - 2026-06-04/05: Reports Meta "repeatedly pushes back" a new AI model release for developers (reported, WSJ via thehindu/dunyanews). - 2026-07-08/09: Meta releases Muse Spark 1.1, unveiled by Zuckerberg; Meta self-reports HLE 58% (with tools/Contemplating mode) vs. independent Artificial Analysis measurement of 39.9% (confirmed release; self-reported score disputed). - 2026-07-09: Commentators (Digg) flag conflict-of-interest concerns over Meta's self-reported Muse Spark 1.1 HLE benchmark claims (reported/opinion). - 2026-08-11/16: Muse Glimmer (open model) and broader Meta AI product rollout continues; no new HLE record claimed (reported). - 2026-08-22: Third-party leaderboard (Artificial Analysis) shows Claude Fable 5 leading HLE at 55.5%, ahead of all Meta figures on standardized measurement (confirmed per artificialanalysis.ai/pricepertoken.com). # Event Will any Meta-published model reach ≥55% HLE accuracy (per agi.safe.ai leaderboard) by Dec 31, 2026? # Outcomes to forecast - Yes (≥55%) - No (<55%) # Kalshi market anchor Kalshi-direct price was NOT returned in this research pass (tool output missing). The closest available cross-market proxy is Polymarket, pricing this same underlying question at **49%** (7d: +2.5%, 30d: +10%, range 28-67.5% over 33 days, $15.3K volume) — trending upward but still near a coin flip. Treat Kalshi price as unknown/gap; Polymarket 49% is the best current consensus proxy. # Sub-question answers 1. **Polymarket price for this/sibling thresholds?** Only this exact 55% threshold market was found (49%); no sibling 25%/40%/70% Meta-HLE markets identified on Polymarket [polymarket_direct/related]. 2. **Highest Meta HLE score on agi.safe.ai to date?** Not directly confirmed on agi.safe.ai itself; via comparable leaderboards, Meta's Muse Spark self-reports 42.8% (no tools)/50.4% (with tools) and a disputed "58%" Contemplating-mode figure; independent Artificial Analysis measures only 39.9%; Scale AI's official leaderboard lists Meta's "Muse Spark" at 40.56% [claude_news, venturebeat, Wikipedia]. 3. **Current cross-lab SOTA and its rise?** As of Aug 2026, Artificial Analysis leaderboard leader is Claude Fable 5 at 55.5%, followed by Claude Opus 5 variants (54.9%, 54.4%); Scale AI's stricter official leaderboard shows lower figures (Gemini 3.1 Pro 46.44%, GPT-5.4 Pro 44.32%). SOTA has risen from ~3% (Jan 2025) to ~41.5% (late 2025) to ~50-55% (mid-2026) [claude_news, code_execution extrapolation]. 4. **Has Meta released a 2026 frontier model with HLE results?** Yes — Muse Spark (April 2026) and Muse Spark 1.1 (July 2026), both from Meta Superintelligence Labs, with published (self-reported) HLE scores [claude_news, gdelt_news, Wikipedia]. 5. **Does Meta publish HLE evals / is it on agi.safe.ai?** Meta publishes its own benchmark claims (self-reported), but independent leaderboards (Artificial Analysis, Scale AI) separately measure and generally show materially lower results than Meta's self-reports; direct confirmation of agi.safe.ai inclusion not found in research, but Meta models appear on comparable/derivative leaderboards [claude_news]. 6. **Shift to closed-source and timeline?** Confirmed — Meta abandoned Llama's open-weight strategy for Muse Spark (proprietary), amid leadership turmoil (Chris Cox exit) and repeated release delays reported through 2025-2026 [cnbc, digitimes, thehindu]. # Key facts (high-confidence, factual) 1. [Wikipedia] Meta Superintelligence Labs released Muse Spark in April 2026, replacing Llama. 2. [venturebeat/deeplearning.ai] Meta self-reports 58% HLE (Contemplating/tool-use mode); independent Artificial Analysis measured only 39.9%. 3. [artificialanalysis.ai via claude_news] Current independently-verified HLE leader (Aug 2026) is Claude Fable 5 at 55.5%, not Meta. 4. [Scale AI/Wikipedia] Official Scale AI leaderboard places Meta's Muse Spark at 40.56%, below GPT-5.4 Pro and Gemini 3.1 Pro. 5. [digitimes/cnbc] Meta delayed its flagship model release and underwent leadership reorg (Cox exit) in late 2025. # Cross-market signals - Kalshi related: No direct sibling HLE markets found; only tangential Meta business markets (headcount, DAP) unrelated to benchmark performance. - Polymarket: 49% YES on this identical question, up from 28% low and trending toward 67.5% high over the past month — indicating rising but contested belief. - Sportsbook implied: N/A. # Analyst opinions and speculation - Commentators (Digg, July 2026) allege conflict-of-interest concerns over Meta's self-reported Muse Spark 1.1 HLE claims, suggesting the 58% figure is not credible/standardized. - Code-execution model estimates ~15-20% probability of YES, citing Meta's persistent historical lag versus frontier labs as the binding constraint, despite ~67% likelihood frontier SOTA overall exceeds 55% by Dec 2026. - Multiple outlets frame Muse Spark as competitive but not frontier-leading (trailing on coding/agentic benchmarks too). # Directional lean per outcome - **Yes**: Meta's own claimed 58% figure exists; Polymarket pricing near 49% shows real market uncertainty, not clear NO consensus; Meta is actively iterating (Spark→1.1) and could improve before Dec 2026. - **No**: All independent/standardized measurements (Artificial Analysis, Scale AI) show Meta below 55% (39.9%-50.4%); resolution source (agi.safe.ai) likely uses standardized methodology, not self-reported multi-agent scores; Meta has a consistent history of lagging frontier by wide margins; conflict-of-interest concerns undermine Meta's self-reported figures. # Gaps / unknowns - No direct Kalshi YES price was retrieved for this ticker — a critical anchor gap. - Unclear whether agi.safe.ai leaderboard would recognize Meta's "Contemplating" tool-augmented mode as a valid comparable HLE Accuracy entry, which is decisive for resolution. - No confirmation of Meta's actual listed score directly from agi.safe.ai (research relied on secondary trackers with conflicting numbers). - Unknown whether Meta will release a further/larger model in Q4 2026 before close. # Calibration anchors - Polymarket proxy price (best available): 49% YES, uptrending (+10% in 30 days). - Independent benchmark data suggests true current Meta score sits ~15pp below threshold on standardized metrics; historical Meta HLE performance has consistently lagged SOTA by large margins (Llama 4: 5-8% vs. contemporaneous SOTA ~20-25%). - Precedent: benchmark self-reports from labs (all vendors) often diverge substantially from third-party verification, typically favoring the vendor.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.32
Yes 35%
No 65%
All independent/standardized measurements place Meta's best 2026 model well below the 55% threshold — Artificial Analysis measured Muse Spark 1.1 at 39.9% and Scale AI's official leaderboard lists 40.56%, versus Meta's disputed self-reported 58% in a tool-augmented 'Contemplating' mode that a standardized agi.safe.ai-style entry likely would not credit. Frontier SOTA itself only reached ~55.5% (Claude Fable 5) by Aug 2026, and Meta has historically lagged the frontier by wide margins (Llama 4 at 5-8% vs ~20-25% SOTA), compounded by leadership turmoil and repeated release delays. Against that, the Polymarket proxy sits at 49% and is uptrending (+10pp in 30 days), reflecting genuine uncertainty about whether Meta's self-reported/tool-mode figure counts and whether a Q4 2026 model lands above threshold. I weight the standardized-measurement evidence and the code-execution estimate (15-20%) more heavily than the market proxy, but not fully, given the resolution-source ambiguity — landing at ~35% YES.
gpt-5.6-sol
0.57
Yes 36%
No 64%
The best consensus anchor is Polymarket at 49% YES, since no direct Kalshi price was retrieved. Evidence points toward No because standardized measurements place Meta around 39.9%-50.4%, while the agi.safe.ai resolution is unlikely to accept Meta's disputed 58% tool-augmented self-report without comparable verification. The broader frontier has reached 55.5%, showing the threshold is attainable, and Meta's rapid Spark-to-Spark-1.1 iteration leaves a meaningful chance of another improvement before year-end. I move below the 49% anchor because Meta has historically lagged frontier labs, suffered release delays, and has not been directly confirmed at or above 55% on the specified leaderboard.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters converge to nearly identical ~35% despite claiming different weighting schemes (one anchors more on standardized data, one on Polymarket) — this convergence masks that neither seriously grapples with why Polymarket sits at 49% and trending toward 67.5% high; if the market has information about an upcoming Q4 model or resolution-criteria clarification, both may be under-weighting that momentum. 2. Neither forecast adequately addresses the critical resolution ambiguity flagged in the brief: "Unclear whether agi.safe.ai leaderboard would recognize Meta's 'Contemplating' tool-augmented mode as a valid comparable HLE Accuracy entry, which is decisive for resolution" — this is treated as a minor haircut rather than potentially the single most decisive factor, and both should have produced conditional estimates (P(Yes | tool-mode counts) vs P(Yes | it doesn't)). 3. Both underweight the "Unknown whether Meta will release a further/larger model in Q4 2026 before close" — given Meta shipped Spark then Spark 1.1 within three months (April→July) and is under competitive pressure post-reorg, a plausible Spark 2.0 or "Avocado 2" release by Q4 could shift standardized scores meaningfully; this base-rate of rapid iteration deserves more explicit modeling than a passing mention. 4. Both cite the code-execution estimate of "15-20% probability of YES" but then land at 35-36%, effectively overriding that estimate by nearly 2x without fully justifying why the Polymarket anchor (a thin, $15.3K-volume, single-market proxy) should dominate a more analytically-grounded estimate — this blending is asserted rather than derived. 5. Neither forecast flags the self-reinforcing measurement problem: since three different trackers (Meta self-report 58%, Artificial Analysis 39.9%, Scale AI 40.56%) disagree by ~18 points on the *same* model, the eventual resolution outcome may hinge on which specific tracker/methodology agi.safe.ai's own leaderboard uses — a factor with binary, not probabilistic, implications that neither analysis models as a discrete branching scenario. 6. The Polymarket price trend (28% low to 67.5% high over 33 days, now 49%, "+10% in 30 days") shows extreme volatility suggesting thin/noisy trading rather than settled consensus; both forecasters treat 49% as a stable anchor rather than flagging the wide range as evidence the market itself doesn't know how resolution will be interpreted, which should widen uncertainty rather than tightly cluster both forecasts at 35-36%.
6. Aggregator Opus · final calibrated forecast
40% (44% confidence)
Yes: 40% No: 60%
Every independently standardized measurement of Meta's best 2026 models sits well below the bar (Artificial Analysis 39.9%, Scale AI 40.56%, Meta's own no-tools figure 42.8%), and Meta has a consistent history of trailing the frontier, which is itself only just at ~55.5%. The main YES pathways are (a) agi.safe.ai crediting Meta's self-reported 58% 'Contemplating' tool-augmented entry — the official HLE leaderboard has historically included tool/search-augmented rows, which makes this a live, quasi-binary branch — and (b) a Q4 Meta release (Spark 2.0) clearing 55% on standardized eval given Meta's rapid April→July iteration cadence. Because these two pathways are more than a minor haircut, I nudge above both forecasters' ~35% toward the Polymarket proxy (49%, uptrending), while staying below it since that market is thin, volatile, and the analytically grounded estimate points lower. Net: ~40% YES, reflecting genuine resolution-criteria ambiguity rather than a settled NO.
Pipeline Timing
Total pipeline time: 163.3s
Per-tool research timings shown in the Research section above.