← Back to scans

Will the highest score achieved by a Meta model on Humanity’s Last Exam in 2026 be 55% or higher?

0xb26b135b38bcbe4c7abe33467a8da10030286d29c4f662ca269ed3ae05db9cbb · Companies · 2026-09-02
20%
Agent
44%
Market Price
-24.0%
Edge
65%
Confidence
Volume: 22,889
Spread: 4.0c
Days to resolution: 120
Markets in event: 4
Final Rationale
The critique correctly identifies that both forecasters anchored too heavily on the 44% market price despite fundamentals pointing much lower: the official leaderboard shows Meta at 40.56%, independent scores have plateaued at 39.9%→45%→44% across three releases, Behemoth is dead, and no frontier release is confirmed before year-end. Closing a 10-15pt gap in four months contradicts Meta's demonstrated trajectory and its pattern of 10-18pt marketing-vs-verified discrepancies. The main residual Yes path is the bimodal rules question — whether ensemble 'Contemplating mode' scores could be credited officially — but the fact that Scale AI already lists Muse Spark at 40.56% (not Meta's claimed 50-58%) suggests the official arbiter has implicitly rejected that methodology, and conflict-of-interest allegations further reduce the chance of a borderline claim being credited. I move below both forecasts to 20% Yes, retaining probability mainly for the eligibility-ruling tail and a possible late-year release, while the thin market ($22.9K volume, declining from 67.5%) deserves limited anchoring weight.
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-24 40% 49% 44%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news claude_news gdelt_news wikipedia
Sub-questions (Fermi decomposition)
  1. What is the current highest HLE accuracy achieved by any model on the official agi.safe.ai leaderboard, and how fast have top scores improved over the past 12 months?
  2. What is the current best HLE score achieved by any Meta model (e.g., Llama 4 variants), and does Meta even appear on the official leaderboard?
  3. What frontier models does Meta (Meta Superintelligence Labs) plan to release in 2026 (e.g., Llama 4 Behemoth, Llama 5), and what benchmark performance is expected or claimed?
  4. Does the official HLE leaderboard include tool-augmented / agentic configurations (which score much higher), and would a Meta model in such a configuration count for resolution?
  5. How far behind OpenAI, Google, Anthropic, and xAI is Meta on reasoning benchmarks generally, and what is the base rate for Meta closing such a gap within one year?
  6. What are related prediction markets pricing for other HLE thresholds (e.g., Meta at 40%, 45%, 50%, 60%) and for other labs' HLE scores in 2026, and are the implied probabilities internally consistent?
Planner reasoning
This is a Polymarket question about whether any Meta model reaches ≥55% HLE accuracy on the official leaderboard by end of 2026. Key drivers are the current market price, the current state-of-the-art on HLE (whether any lab has reached 55%), Meta's current best HLE score and its 2026 model roadmap (Llama 4 Behemoth, Meta Superintelligence Labs), and whether leaderboard scores count tool-use configurations. I'll anchor on the direct market, triangulate with related markets on other thresholds and labs, and use news searches for leaderboard and Meta model developments.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.4s 1 ## This Market's Polymarket Data **Will the highest score achieved by a Meta model on Humanity’s Last Exam in 2026 be 55% or higher?** - Current price (probability): 44.00% - 7-day price change: -5.50% - 30-day price change: -2.00% - Total volume: $22,889 (USD notional) - Price range: 28.00% - 67.5
polymarket_related OK 3.4s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword "Humanity's Last Exam": 0 markets | keyword 'HLE': 0 markets | keyword 'Meta model': 0 markets | keyword 'Llama': 0 markets
kalshi_related OK 3.3s 3 3 related markets / summaries. keyword "Humanity's Last Exam": ok | keyword 'Meta AI': ok | keyword 'Llama benchmark': ok
claude_news OK 25.2s 7 Key findings on HLE leaderboard and Meta's standing (as of Sept 2026): - **Overall leaderboard top scores are now in the 55–65% range**, driven mainly by Anthropic. Per BenchLM, "Claude Fable 5.1 leads the HLE leaderboard with 65%, followed by Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%)" as
claude_news OK 26.2s 14 **Key findings on Meta's 2026 frontier AI models and HLE performance:** - **Llama 4 Behemoth is effectively dead.** Behemoth is effectively shelved — never publicly released, never formally cancelled; mid-training MoE-routing and chunked-attention issues at 2T-scale meant Meta lost confidence the
gdelt_news OK 91.2s 10 GDELT: 10 articles across 3 queries (lookback=60d). "Meta Llama Humanity's Last Exam": 10 hits | 'Meta Superintelligence Labs model release': error HTTPSConnectionPool(host='api.gdeltproject.org', port=443): Max retries exceeded with url: /api/v2/doc/doc?query=%28meta+OR+superintelligence+OR+labs+OR
wikipedia OK 0.2s 3 Fetched 3 Wikipedia entries (0 missing pages).
3. Evidence Brief Sonnet · 7575 chars
# Current state The market resolves on the official agi.safe.ai/Scale AI HLE leaderboard, where the current best documented Meta entry (Muse Spark, MSL's post-Llama flagship) sits at ~40-45% per independent trackers — well below the 55% bar. Meta's own marketing claims of 50.2%-58% rely on an ensemble "Contemplating mode" not obviously equivalent to the standard "HLE Accuracy" leaderboard metric, and are contested by independent evaluators (Artificial Analysis) and press (conflict-of-interest allegations). # Timeline of key events - 2025-04: Meta releases Llama 4 (Scout/Maverick); best documented HLE score ~6.2% (rank ~98/97) — confirmed low benchmark, per LayerLens/BenchmarkList (claude_news). - Late 2025–early 2026: Llama 4 Behemoth (2T param) finishes training but is shelved due to "poor internal performance"; never formally released or cancelled (reported, NYT via claude_news). - 2026-04-08: Meta Superintelligence Labs (led by Alexandr Wang) launches Muse Spark, a closed-weight model replacing the Llama line (confirmed, Wikipedia + officechai.com + claude_news). - 2026-04: Muse Spark 1.0 launch — Meta self-reports 50.2% HLE (no tools, "Contemplating" mode) and a separate 58% figure for full Contemplating mode; independent Artificial Analysis measures Thinking-mode score at 39.9% (reported/disputed). - 2026-07-09: Commentators (Digg) allege conflict-of-interest issues compromising Meta's Muse Spark 1.1 HLE benchmark claims (reported). - 2026-07 (undated, ~mid): Muse Spark 1.1 released; independent HLE ≈45% (Artificial Analysis), Intelligence Index 51 (reported). - 2026-08-11: Meta unveils Muse Glimmer, a smaller 30B local-run model (not a frontier/HLE contender) (confirmed, Forbes). - 2026-08-16: Official Scale AI public leaderboard snapshot: Gemini 3.1 Pro (thinking high) 46.44%, GPT-5.4 Pro 44.32%, Muse Spark 40.56%, Gemini 3 Pro 37.52% (confirmed per IntuitionLabs summary of agi.safe.ai partner site). - 2026-08 (~"Muse Spark 1.2"): Independent HLE score dips to ~44% (from 45%), Intelligence Index rises to 54 (reported, Artificial Analysis). - 2026-09-02: Frontier field leaders (non-Meta) hit 55-65% on various trackers — Claude Fable 5.1 at 65% (BenchLM) / 59.1% (Artificial Analysis) — confirming the 55%+ bar is achievable technically but Meta remains ~15-20 points behind (confirmed). # Event Will any Meta-published model reach ≥55% accuracy on the official Humanity's Last Exam (agi.safe.ai) leaderboard by Dec 31, 2026? # Outcomes to forecast - Yes (Meta model hits ≥55% HLE) - No (Meta stays below 55%) # Kalshi market anchor Current YES price: **44%** (Polymarket data, same ticker). Trend: -5.5% over 7 days, -2% over 30 days — declining. Range has been wide (28%–67.5%) over 42 days of trading, volume modest (~$22.9K total). Market has cooled from earlier optimism, consistent with underwhelming independent Muse Spark verification scores. # Sub-question answers 1. **Current top HLE score / 12-month trend** — Independent trackers show frontier leaders now at 55-65% (Claude Fable 5.1: 65% per BenchLM, 59.1% per Artificial Analysis, Sept 2026), up from single digits in early 2025 (GPT-4o 2.7%, Claude 3.5 Sonnet 4.1%) — an unprecedented ~1-year climb led by Anthropic/OpenAI/Google. The official Scale AI leaderboard is lower/lagged, topping out near 46% (Gemini 3.1 Pro, Aug 2026). 2. **Best Meta score / leaderboard presence** — Meta's Muse Spark appears on the official Scale AI leaderboard at 40.56% (Aug 2026, IntuitionLabs). Independent AA trackers put Muse Spark 1.1/1.2 at 44-45%. Meta's own marketing claims 50.2-58% via a multi-agent "Contemplating mode," disputed by independent evaluators and press (conflict-of-interest allegations, Digg). 3. **2026 frontier model plans** — Llama 4 Behemoth is effectively dead/shelved (never shipped). Meta pivoted fully to closed-weight Muse Spark line under MSL (launched April 2026), with iterative updates (1.0→1.1→1.2) showing incremental, sometimes regressing, gains (Intelligence Index 43→51→54). Muse Glimmer (Aug 2026) is a smaller local model, not a benchmark contender. 4. **Tool-augmented/agentic configs** — Meta's highest self-reported figures (58%) come from an ensemble "Contemplating mode," which functions similarly to tool/agentic augmentation; independent verifiers don't replicate this in single-model "Thinking mode" tests (39.9-45%). Whether the official leaderboard would accept such a mode as "HLE Accuracy" is unclear/disputed — a key resolution risk. 5. **Gap vs. OpenAI/Google/Anthropic/xAI** — Meta trails significantly: on the official leaderboard, Meta (40.56%) sits below GPT-5.4 Pro (44.32%) and Gemini 3.1 Pro (46.44%), and far below Anthropic's ~60-65% frontier scores. Meta also lags on coding/hard reasoning generally (Llama 4 "trails DeepSeek V4 and Qwen 3.5/3.6"). No base rate exists for Meta closing a 15-20pt benchmark gap in under 4 months given repeated underperformance vs. self-reported claims. 6. **Related markets pricing** — No other HLE-threshold or lab-specific Kalshi/Polymarket markets found (0 matches on Polymarket related-market scan); this ticker is the only relevant HLE market identified. # Key facts (high-confidence, factual) 1. [Polymarket] Current YES price 44%, declining trend. 2. [IntuitionLabs/agi.safe.ai] Official leaderboard: Muse Spark at 40.56% (Aug 16, 2026); top model (Gemini 3.1 Pro) at 46.44%. 3. [Wikipedia] Muse Spark replaced Llama as Meta's flagship in April 2026 under MSL. 4. [claude_news/NYT] Behemoth shelved, never released. 5. [Artificial Analysis via claude_news] Independent Muse Spark scores: 1.0 39.9%, 1.1 ~45%, 1.2 ~44% — plateauing below 55%, not converging upward. # Cross-market signals - Kalshi related: No direct HLE markets found besides this ticker; tangential Meta markets (headcount, antitrust) unrelated to AI benchmarks. - Polymarket: Same ticker mirrored — 44% YES, high volatility (28-67.5% range), suggesting past optimism (possibly around Meta's 58% self-reported claim) has since faded as independent verification came in lower. - Sportsbook implied: N/A. # Analyst opinions and speculation - Claude_news synthesis: reaching 55%+ "would require an enormous, unprecedented leap...inconsistent with Meta's current trajectory." - Digg commentary alleges conflict-of-interest in Meta's self-reported benchmark methodology, casting doubt on Meta's higher claimed figures being accepted for official resolution. # Directional lean per outcome - **Yes**: Field-wide progress is rapid (frontier leaders already exceed 55-65%); Meta's self-reported "Contemplating mode" claims 58%, and if such ensemble scoring were credited on the official leaderboard, Yes could resolve. Four more months of releases remain. - **No** (favored): Official leaderboard has Meta at ~40.56%; independent trackers show a plateau (~44-45%) across three iterative releases with no clear upward trajectory; flagship Behemoth is dead; Meta trails all four major competitors by double digits. # Gaps / unknowns - Whether "Contemplating mode" ensemble scores would count as official "HLE Accuracy" per resolution rules is unresolved and pivotal. - No confirmation of further Meta frontier releases planned before Dec 31, 2026. - Official agi.safe.ai leaderboard update cadence/lag vs. third-party trackers unclear. # Calibration anchors - Kalshi/Polymarket current YES price: 44% (anchor), down from a 67.5% high. - Precedent: Meta's Muse Spark line has repeatedly underperformed its own marketing claims by 10-18 points when independently verified (1.0: 58% claimed vs 39.9% verified).
4. Forecasters Opus + GPT · in parallel
claude-fable-5
0.62
Yes 28%
No 72%
The official agi.safe.ai leaderboard has Meta's Muse Spark at 40.56% (Aug 2026), and independent Artificial Analysis scores show a plateau at 39.9%→45%→44% across three releases — no upward trajectory toward 55%. Meta's 58% self-reported 'Contemplating mode' claim has been contested by independent evaluators and press conflict-of-interest allegations, and it's unclear it would count on the official leaderboard. The Kalshi/Polymarket anchor is 44% YES but declining sharply (-5.5% over 7 days, down from 67.5%), and the fundamentals (dead Behemoth, 15-20pt gap to Anthropic, repeated marketing-vs-verified discrepancies of 10-18pts) suggest the market still overweights Meta's self-claims. I discount below the anchor to ~28%, retaining meaningful Yes probability for a surprise Q4 release or the ensemble mode being credited officially.
gpt-5.6-sol
0.68
Yes 36%
No 64%
The 44% YES market price is the anchor, but Meta’s official score of 40.56% and independent scores near 44–45% imply a substantial gap to 55%. Muse Spark’s successive versions have plateaued or regressed slightly, while Behemoth was shelved and no stronger 2026 release is confirmed, pointing toward No despite rapid field-wide progress. Meta’s claimed 58% Contemplating result preserves a meaningful Yes path, but its methodological eligibility is unclear and prior Meta claims have exceeded independent results by 10–18 points. I therefore move below the market consensus while retaining substantial probability for a late release or acceptance of an ensemble configuration.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasts land materially above what the plateau data implies (three consecutive Meta releases at 39.9%→45%→44%, official leaderboard 40.56%) — a ~10-15pt gap to close in ~4 months with a "dead" flagship (Behemoth) and no confirmed pipeline; a base-rate/momentum-only read would justify Yes well under 20%, yet both sit at 28-36%, seemingly anchored too heavily on the Kalshi price rather than the fundamentals. 2. Neither forecaster grapples with the resolution-criteria ambiguity as a bimodal risk rather than a continuous discount: if "Contemplating mode" is ruled ineligible (likely, given it's an ensemble/agentic scheme not matching standard "HLE Accuracy"), Yes probability should collapse near-zero; if ruled eligible, Meta's self-reported 58% would trivially satisfy the bar. Averaging into a smooth 28-36% may misrepresent this as continuous uncertainty rather than a discrete rules-resolution gate. 3. Both cite the declining market trend (67.5%→44%, -5.5%/7d) as reinforcing their downward adjustment, but neither asks why the market still sits as high as 44% given such weak fundamentals — this could indicate residual uncertainty specifically about the ensemble-mode rules question (sub-question 4), which deserves more explicit weight rather than being folded into a generic "discount below anchor." 4. Neither forecast addresses potential leaderboard lag/update-cadence risk (flagged in Gaps/Unknowns): if the official agi.safe.ai leaderboard updates infrequently, a hypothetical Meta breakthrough late in Q4 2026 might not be reflected/certified by the Dec 31 close, which should push probability further toward No than either forecast allows. 5. Both retain "meaningful" Yes probability for a "surprise Q4 release," but the evidence brief offers no signal of any planned Meta frontier release beyond Muse Spark iterations and Glimmer (a smaller, non-frontier model) — this hedge appears unanchored to any concrete evidence and may be inflating Yes beyond what the brief supports. 6. Neither forecaster explicitly weighs the conflict-of-interest allegations (Digg, July 2026) as bearing on resolution risk — if Meta's benchmark methodology is viewed as compromised, official arbiters (Scale AI) may be less likely to credit any borderline high score, which is a one-directional risk toward No that isn't reflected in either rationale.
6. Aggregator Opus · final calibrated forecast
20% (65% confidence)
Yes: 20% No: 80%
The critique correctly identifies that both forecasters anchored too heavily on the 44% market price despite fundamentals pointing much lower: the official leaderboard shows Meta at 40.56%, independent scores have plateaued at 39.9%→45%→44% across three releases, Behemoth is dead, and no frontier release is confirmed before year-end. Closing a 10-15pt gap in four months contradicts Meta's demonstrated trajectory and its pattern of 10-18pt marketing-vs-verified discrepancies. The main residual Yes path is the bimodal rules question — whether ensemble 'Contemplating mode' scores could be credited officially — but the fact that Scale AI already lists Muse Spark at 40.56% (not Meta's claimed 50-58%) suggests the official arbiter has implicitly rejected that methodology, and conflict-of-interest allegations further reduce the chance of a borderline claim being credited. I move below both forecasts to 20% Yes, retaining probability mainly for the eligibility-ruling tail and a possible late-year release, while the thin market ($22.9K volume, declining from 67.5%) deserves limited anchoring weight.
Pipeline Timing
Total pipeline time: 187.7s
Per-tool research timings shown in the Research section above.