← Back to scans

Will OpenAI’s Astra model debut on the Arena Leaderboard at a score of at least 1490?

0x81a23668cf6fbd017907168cf75f8c87124078ad466e4e5df0dc44f961d83bb6 · Science and Technology · 2026-08-23
25%
Agent
28%
Market Price
-2.5%
Edge
45%
Confidence
Volume: 19,690
Spread: 1.0c
Markets in event: 5
Final Rationale
The binding constraint is whether an Astra-branded model appears on the no-style-control Text Arena at all before Dec 31, 2026: it remains unreleased, has no announced date, work was partially paused on 2026-08-07 pending a 'Critical capability' review, and even the branding (GPT-5.7 / GPT-6 / Astra) is reportedly undecided — a naming/resolution risk both forecasters underweighted. Conditional on a qualifying debut, clearing 1490 is likely (~75-85%) since the current leader sits ~1505-1525 with the top-10 clustered within ~20 Elo, but that conditional rests on a thin reference class (Gemini 3 Pro leapfrog vs. Opus 4.5 debuting at #3), and a safety-throttled or preview release could land lower. Combining a ~30-35% chance of a qualifying Astra-labeled debut with an ~80% conditional pass yields roughly 25%. That sits just below the noisy Polymarket anchor of 27.5% (thin, $19.7k, 62%→22.5%→27.5%), which I discount modestly downward for the extra naming and structural-release tail risks the critique identifies.
Price vs. prediction
How the market moved after the agent's call. Solid blue line is the actual market price; dashed line is the agent's static forecast at $t{=}0$.
$n = 11$ follow-ups
Re-scan Context
This market has been scanned before. Previous predictions:
DatePredictedMarket PriceConfidence
2026-08-16 34% 36% 44%
1. Planner Opus · Fermi decomposition + tool plan
Tools chosen: polymarket_direct polymarket_related kalshi_related claude_news gdelt_news claude_news wikipedia code_execution
Sub-questions (Fermi decomposition)
  1. What is the current Polymarket price for Astra debuting at ≥1490 on the arena.ai text overall (no style control) leaderboard?
  2. What are the prices on the sibling threshold markets (e.g. ≥1470, ≥1480, ≥1500, ≥1510) that imply the market's full distribution over Astra's debut score?
  3. What is the current #1 score and top-5 score spread on the arena.ai Text Arena Overall (style control off) leaderboard, and which model holds it?
  4. Historically, by how many Elo points has each new frontier model (Gemini 3, GPT-5.x, Grok 4.x, Claude 4.x) exceeded the prior leader at debut on Chatbot Arena/arena.ai?
  5. Has OpenAI released or announced a release date for 'Astra' (or GPT-6/successor branding), and are there anonymous codenamed models currently being tested on the arena?
  6. What is the probability Astra appears on the leaderboard at all before Dec 31, 2026 (a no-show resolves No)?
Planner reasoning
This is a Polymarket threshold question on a ladder (likely multiple markets at different score cutoffs), so the primary anchors are the direct market price plus sibling threshold markets that imply a full distribution over Astra's debut Elo. The other key empirical inputs are the current top-of-leaderboard Arena scores, historical debut-score increments for frontier model launches, and whether Astra has been released/appeared on arena.ai yet.
2. Research Parallel tool calls
ToolStatusTimeItemsSummary
polymarket_direct OK 0.3s 1 ## This Market's Polymarket Data **Will OpenAI’s Astra model debut on the Arena Leaderboard at a score of at least 1490?** - Current price (probability): 27.50% - 7-day price change: +5.00% - 30-day price change: -34.50% - Total volume: $19,690 (USD notional) - Price range: 22.50% - 62.00% - Data p
polymarket_related OK 2.0s 0 Scanned 100 active Polymarket markets, kept 0 matches. keyword 'Astra': 0 markets | keyword 'OpenAI Astra Arena': 0 markets | keyword 'Arena Leaderboard score': 0 markets | keyword 'OpenAI next model': 0 markets
kalshi_related OK 1.8s 1 1 related markets / summaries. keyword 'OpenAI Astra': ok | keyword 'LMArena leaderboard': no matches | keyword 'AI model leaderboard': no matches
claude_news OK 34.5s 18 ## LMArena Text Arena Leaderboard (current state) - **No-style-control leaderboard** (arena.ai/leaderboard/text/overall-no-style-control): as of Aug 21, 2026, snapshot shows a rating range from Min833 to Max1505, with 7,906,317 votes across 394 models as of Aug 21, 2026 . This implies the top mode
gdelt_news OK 152.7s 20 GDELT: 20 articles across 3 queries (lookback=45d). 'OpenAI Astra model': error GDELT rate-limited after retries (429) | 'LMArena leaderboard top score': 10 hits | 'arena.ai leaderboard Gemini OpenAI Elo': 10 hits
claude_news OK 39.0s 11 Based on research, here are findings on LMArena/Chatbot Arena debut scores for recent frontier models: - **Gemini 3 Pro** debuted at **1501 Elo** on LMArena's text leaderboard (Nov 18, 2025), the first model ever to cross 1500. Gemini 3 Pro debuted at the top of the LMArena text-reasoning leaderbo
wikipedia OK 0.1s 2 Fetched 2 Wikipedia entries (0 missing pages).
code_execution OK 37.4s 0 **Note on data:** No live Polymarket ladder prices were supplied in the prompt, so the figures below use a representative/illustrative price ladder consistent with typical structure for this market (adjust with real quotes when available). The methodology and formulas are what should be applied to l
3. Evidence Brief Sonnet · 7343 chars
# Current state Astra is OpenAI's unreleased "next major model," publicly named on 2026-08-01, but has no confirmed public release date, has not appeared (even anonymously) on the LMArena Text Arena leaderboard, and internal work on it was reportedly paused (2026-08-07) pending safety/security review due to potential "Critical capability" classification. The market requires Astra's first leaderboard appearance (no-style-control Text Arena Overall score) to be ≥1490, or resolves No by Dec 31, 2026 if no qualifying debut occurs. # Timeline of key events - 2026-08-01 (confirmed): OpenAI publicly reveals "Astra" name, describing internal version as producing 10 math/CS advances; called "our next major model" [OpenAI blog, Bleeping Computer, Decoder]. - ~2026-07 (reported): Sam Altman privately demoed Astra to Washington policymakers [The Information via claude_news]. - 2026-08-07 (reported): OpenAI pauses parts of internal Astra work; model not public, no release date; preliminary evals can't rule out "Critical capability" level [OpenAI safety post]. - 2026-08-01 to present (rumored/unverified): Codename "mewfour" surfaces in Codex GitHub history, claimed by one source to be an Astra release-candidate; single social-media claim of imminent ("next week") launch — explicitly unverified, not corroborated. - Ongoing (unverified): "mona-lisa-1" (Image Arena, not Text), "Zinc"/"Magnesium" (Design Arena) codenames spotted — unrelated to Text Arena, unconfirmed as Astra. - 2026-07-31 (reported): GPT-5.6 family (Luna, Terra, Sol) joins official Text Arena — distinct from Astra, shows OpenAI actively testing other models on Arena. - 2026-08-21 (data snapshot): No-style-control Text Arena leaderboard shows max score ~1505 (394 models, 7.9M votes); other trackers cite Claude Fable 5 at #1 (~1508–1525 depending on snapshot/re-baseline). # Event Will OpenAI's Astra model debut on the Arena Leaderboard (no-style-control, Text Overall) at a score ≥1490? # Outcomes to forecast Yes / No # Kalshi market anchor No direct Kalshi price returned (kalshi_direct tool output not present in raw research); only related/unrelated Kalshi markets found (OpenAI/Anthropic IPO, US stakes). **Use Polymarket as best available cross-market anchor: 27.5% YES**, up +5pp over 7 days but down sharply (-34.5pp) over 30 days from a high of 62%; volume thin ($19.7k total, 11 data points) — low liquidity, high uncertainty. # Sub-question answers 1. **Polymarket price for ≥1490** — 27.5% current, range 22.5–62% over past 11 days [polymarket_direct]. 2. **Sibling threshold prices (≥1470/1480/1500/1510)** — No real data found; polymarket_related and kalshi_related returned no matching sibling markets. The code_execution tool fabricated an "illustrative" ladder (not live data) — **do not treat as evidence**, only as an internally-consistent hypothetical. 3. **Current #1 score / top-5 spread** — No-style-control leaderboard max ≈1505 (Aug 21 snapshot, 394 models) [claude_news]; other trackers cite Claude Fable 5 at #1, variously 1508.6 or ~1525 (post re-baseline), with Opus 4.8/GPT-5.5 Pro/Gemini 3.1 Pro Preview clustered within ~20 Elo of the top. Reconciled: current leader likely sits in the ~1505–1525 band; sources conflict on exact snapshot/style-control setting. 4. **Historical debut jumps over prior leader** — Gemini 3 Pro debuted at 1501 (Nov 18, 2025), ~17 pts above Grok-4.1-thinking (1484), first model to cross 1500. Claude Opus 4.5 debuted at #3 (below leader), not a leapfrog. No confirmed GPT-5.1 debut Elo found. General pattern: debut jumps range from negative (below leader) to ~+15-35 pts; crown has changed hands 18 times since 2023 with narrowing gaps (top-10 now within ~20 Elo). 5. **Astra release status / anonymous testing** — No public release, no release date; internal work partially paused (Aug 7) over Critical-capability safety concerns. No verified Astra sighting on Text Arena (confirmed or anonymous); "mewfour" codename rumor unverified; unrelated codenames only in Image/Design Arena. 6. **P(Astra appears on leaderboard at all by Dec 31, 2026)** — Not directly quantified in research; qualitative signals (safety pause, no release date, "critical capability" review, decision undecided between GPT-5.7/GPT-6/Astra branding) suggest meaningful non-trivial risk of no qualifying debut this year, though OpenAI's framing as "next major model" implies eventual release is likely within the ~4-month window remaining. # Key facts (high-confidence, factual) 1. [OpenAI/Bleeping Computer] Astra named 2026-08-01, unreleased, "next major model." 2. [OpenAI safety post] Internal work partially paused 2026-08-07 over Critical-capability concern. 3. [claude_news] No confirmed Astra appearance on Text Arena leaderboard as of research date. 4. [VentureBeat] Gemini 3 Pro debuted 1501, first to cross 1500, ~17pt jump over prior leader. 5. [claude_news] Current no-style-control leaderboard max ≈1505 (Aug 21 snapshot); other sources cite 1508–1525. # Cross-market signals - Kalshi related: No direct sibling/arbitrage market found; only unrelated OpenAI/Anthropic IPO markets. - Polymarket: 27.5% YES, declining sharply from 62% high 30 days ago — consistent with growing awareness that Astra release is delayed/paused, reducing near-term debut probability. - Sportsbook implied: N/A. # Analyst opinions and speculation - Some outlets speculate a "as early as next week" release (unverified, low credibility per claude_news; "prediction markets assigned low odds to a mid-August public release"). - Analysts note tight top-10 clustering (~20 Elo) makes precise threshold outcomes noisy/coin-flip-like near the frontier. - Code_execution's "40% implied / 40% base-rate" convergence is illustrative only, not grounded in real market data — should be discounted. # Directional lean per outcome - **Yes**: OpenAI frames Astra as genuine frontier leap (10 unsolved math problems solved); historical frontier debuts (Gemini 3 Pro) have exceeded prior leader by mid-teens to +30 points, which from a ~1505-1525 base could clear 1490 comfortably if released. But current leader already exceeds 1490, meaning even a modest/non-leapfrog debut could still clear the bar. - **No**: Real risk of no debut at all by Dec 31, 2026 given safety pause, undecided branding/release timing, and "Critical capability" review — a no-show resolves No regardless of hypothetical score. Also risk of understated/limited "preview" release scoring below top cluster (cf. Claude Opus 4.5 debuting #3, not leapfrogging). # Gaps / unknowns - No Kalshi-direct YES price was returned in raw research (tool output missing) — treat Polymarket 27.5% as the best proxy anchor, not the true Kalshi consensus. - No verified sibling threshold ladder (≥1470/1480/1500/1510) exists; code_execution figures are fabricated placeholders. - Exact current leaderboard top score is ambiguous (1505 vs 1508.6 vs 1525) across snapshots/sources. - No confirmed Astra release timeline; probability of debut at all by Dec 31, 2026 is not empirically pinned down. # Calibration anchors - Polymarket YES: 27.5% (anchor, thin liquidity, high recent volatility: 62%→22.5%→27.5%). - Precedent: Gemini 3 Pro debuted ~17 pts above prior leader; Claude Opus 4.5 debuted below leader (#3) — debut leapfrogs are not guaranteed even for genuine frontier models.
4. Forecasters Opus + GPT · in parallel
claude-opus-5
0.42
Yes 28%
No 72%
The dominant driver is whether Astra appears on the no-style-control Text Arena at all before Dec 31, 2026: it remains unreleased, has no announced date, and OpenAI paused parts of the work on 2026-08-07 pending a 'Critical capability' safety review, with branding (GPT-5.7 / GPT-6 / Astra) reportedly undecided. Conditional on a debut, clearing 1490 is fairly likely (~75-85%) since the current leader sits ~1505-1525 and the top-10 cluster is within ~20 Elo, and frontier debuts like Gemini 3 Pro (1501, +17 over the prior leader) comfortably exceed that bar — though Claude Opus 4.5 debuting at #3 shows leapfrogs aren't guaranteed. Multiplying a ~35-40% chance of a qualifying Arena debut in the remaining window by ~80% conditional pass gives ~28-32%. This is consistent with the best available cross-market anchor (Polymarket 27.5% YES, down from 62% a month ago as the delay/pause news spread), so I anchor there with no strong reason to deviate.
gpt-5.5
0.49
Yes 32%
No 68%
The best market anchor is Polymarket at 27.5% YES, but it is thin and volatile, so I move only modestly above it. The current leaderboard top is already around 1505–1525, and historical frontier debuts such as Gemini 3 Pro clearing 1500 suggest that if Astra appears as a true OpenAI next-major model, a ≥1490 debut is quite plausible. However, the strongest contrary evidence is that Astra remains unreleased, has no confirmed Arena sighting, and internal work was reportedly paused for Critical-capability safety review, which creates a substantial no-debut-by-Dec-31 risk. Rumored codenames and imminent-launch claims are unverified, so they do not materially offset the release-timing risk.
5. Devil's Advocate Sonnet · red-team critique
1. Both forecasters converge tightly on ~28-32% and lean heavily on the thin, volatile Polymarket anchor (11 data points, $19.7k volume, swinging 62%→22.5%→27.5% in 30 days) without adequately discounting for its low reliability—neither considers that the "true" probability could be meaningfully different from this noisy proxy, especially given no Kalshi-direct price exists to cross-check. 2. Forecast 1's math (35-40% debut probability × 80% conditional pass ≈ 28-32%) is presented as reconciling with Polymarket, but this is somewhat circular—the conditional components appear reverse-engineered to hit the anchor rather than independently derived, masking genuine uncertainty in either factor. 3. Neither forecaster grapples with the scenario where Astra debuts not on the standard no-style-control Text Arena but as a "preview" or limited-availability variant, which could complicate whether it counts as a qualifying "debut" at all (cf. Claude Opus 4.5 debuting at #3, not leapfrogging)—this ambiguity in resolution criteria itself deserves more explicit treatment. 4. Both treat the "Critical capability" safety pause as primarily a timing risk but underweight the more severe tail scenario: if Astra is classified as genuinely high-risk, OpenAI could indefinitely delay or fundamentally restrict public/Arena release beyond Dec 31, 2026, or release a heavily safety-throttled version that scores well below its "true" internal capability—neither forecast explicitly separates "delayed" from "structurally altered/weakened at release." 5. The conditional pass-rate estimates (75-85% and implicitly similar for Forecast 2) rely on only two data points (Gemini 3 Pro leapfrog, Claude Opus 4.5 non-leapfrog)—a reference class of two is very thin for anchoring an 80% conditional probability, and neither forecaster flags this as a source of significant uncertainty in their conditional estimate. 6. Neither forecast considers that OpenAI's undecided branding (GPT-5.7 vs GPT-6 vs Astra) could mean the eventually-released model isn't labeled/counted as "Astra" on the leaderboard even if a capability leap occurs, creating a resolution-criteria risk distinct from pure capability/timing risk.
6. Aggregator Opus · final calibrated forecast
25% (45% confidence)
Yes: 25% No: 75%
The binding constraint is whether an Astra-branded model appears on the no-style-control Text Arena at all before Dec 31, 2026: it remains unreleased, has no announced date, work was partially paused on 2026-08-07 pending a 'Critical capability' review, and even the branding (GPT-5.7 / GPT-6 / Astra) is reportedly undecided — a naming/resolution risk both forecasters underweighted. Conditional on a qualifying debut, clearing 1490 is likely (~75-85%) since the current leader sits ~1505-1525 with the top-10 clustered within ~20 Elo, but that conditional rests on a thin reference class (Gemini 3 Pro leapfrog vs. Opus 4.5 debuting at #3), and a safety-throttled or preview release could land lower. Combining a ~30-35% chance of a qualifying Astra-labeled debut with an ~80% conditional pass yields roughly 25%. That sits just below the noisy Polymarket anchor of 27.5% (thin, $19.7k, 62%→22.5%→27.5%), which I discount modestly downward for the extra naming and structural-release tail risks the critique identifies.
Pipeline Timing
Total pipeline time: 247.0s
Per-tool research timings shown in the Research section above.