# Current state
Astra is OpenAI's unreleased, internally-paused "next major model" (disclosed by name Aug 1, 2026); it has no public release date, no API access, and has NOT appeared on the Arena.AI/LMArena leaderboard yet. Resolution requires Astra to (1) actually appear on the leaderboard (non-AutoEval) by Dec 31, 2026 AND (2) debut at a displayed score ≥1490.
# Timeline of key events
- **2026-08-01** (confirmed): OpenAI publicly names "Astra" as its next major model in a blog post ("Ten advances in mathematics"), describing it as a new tier alongside July's Sol/Terra/Luna releases, not a GPT-5.6 patch.
- **2026-08-01/02** (reported, BleepingComputer): Astra reportedly solved 10 long-standing math problems (geometry, coding theory, complexity, cryptography, combinatorics); OpenAI has not decided final branding (GPT-5.7, GPT-6, or "Astra" itself).
- **2026-08-01/02** (reported, The Hacker News): OpenAI's preliminary evals flag Astra may hit "Critical" cyber capability level; internal activities involving Astra are being paused pending enhanced security controls (isolated environments, restricted tool/network access, weight protections).
- **Ongoing, as of research date** (confirmed via absence of evidence): No LMArena/Arena.AI listing for Astra exists; model remains internal-only with no public access.
# Event
Will OpenAI's Astra model debut on the Arena.AI Text Leaderboard (no style control) at a score of at least 1490?
# Outcomes to forecast
- Yes (Astra debuts at ≥1490)
- No (Astra debuts below 1490, is AutoEval-only, doesn't appear by Dec 31 2026, or naming/consensus fails to qualify)
# Kalshi market anchor
No direct Kalshi price returned for this ticker in tool output; only Polymarket data available for the same event. **Polymarket YES price: 34.5%**, up +7pts over 7 days but down -27.5pts over 30 days (range 22.5%–62%, $21K volume, 19 data points) — suggesting the market has cooled substantially from an early-August peak near 62%, likely reflecting the safety-pause news reducing near-term release odds, with a small recent bounce.
# Sub-question answers
1. **Current leaderboard scores** — Estimates vary by tracker: llm-stats.com puts the current #1 (Grok-4.1 Thinking) at 1483; localaimaster.com and swfte.com put Claude Opus 4.8 / Gemini 3.1 Pro / GPT-5.5 Pro cluster at ~1500–1525, with three models above the 1500 Elo barrier. Code-execution analysis anchors current top at ~1501 (Gemini 3 Pro, rescaled). OpenAI's best (GPT-5.5 Pro/GPT-5.6 Sol) sits roughly 1490–1510 per secondary sources — no primary lmarena.ai page was retrieved, so treat as approximate. [claude_news; code_execution]
2. **Astra release/availability** — No confirmed release date; internally paused over cyber-capability safety concerns; branding undecided (could ship as GPT-5.7, GPT-6, or "Astra"); not publicly accessible, so no Arena listing is imminent. [BleepingComputer, TheHackerNews via claude_news]
3. **Historical debut jumps** — Across 9 frontier debuts (Mar'24–Nov'25), jump vs. prior #1 ranged -9 to +51 Elo, mean +26.6, std 20.9 (e.g., Grok-4 debuted -9 below prior top; Gemini 3 Pro's +51 was largest on record). [code_execution]
4. **Sibling/related markets** — No sibling Polymarket threshold markets (≥1450, ≥1470, ≥1500) were found in the scan (0 matches for "Astra," "Arena leaderboard," "OpenAI model," "GPT-6"); only this single market's own price history is available. Kalshi has no Astra-specific market; a related "OpenAI/Anthropic IPO" market exists but is not informative here. [polymarket_related; kalshi_related]
5. **No-style-control vs. style-control scores** — Not directly quantified in research; the market explicitly uses no-style-control scores, and current cluster estimates (1480–1525) are presumed to already reflect that leaderboard variant per trackers cited, though exact style-control deltas weren't isolated in the data. Gap/unknown.
6. **AutoEval / branding risk** — Not directly addressed by research; the resolution rules explicitly exclude AutoEval-labeled entries and require confirmation the model is "Astra" (or a confirmed successor/rebrand). Given OpenAI hasn't settled naming (could ship as GPT-5.7/GPT-6), there's real risk of ambiguity in whether a future release counts as "Astra" under the rules — flagged as unresolved risk. [BleepingComputer]
# Key facts (high-confidence, factual)
1. [OpenAI blog, Aug 1 2026] OpenAI publicly named "Astra" as its next major model in a math-focused announcement.
2. [TheHackerNews] OpenAI has paused internal Astra activities pending enhanced security controls due to potential "Critical" cyber capability.
3. [BleepingComputer] Final branding for the eventual release (Astra vs. GPT-5.7 vs. GPT-6) is undecided.
4. [multiple trackers] No official confirmation Astra has appeared on any Arena leaderboard.
5. [Polymarket] Market priced this event as high as 62% in early August, now down to 34.5%.
# Cross-market signals
- Kalshi related: No Astra-specific Kalshi market found; unrelated OpenAI/Anthropic IPO market at 93% (not informative).
- Polymarket: Same-event market at 34.5% YES, down sharply from 62% peak — market has become more pessimistic, likely driven by safety-pause/no-release-date news rather than doubts about score threshold itself.
- Sportsbook implied: N/A.
# Analyst opinions and speculation
- Code-execution model: conditional P(score≥1490 | Astra appears) ≈ 0.93–00.96 (bar is easy to clear vs. current ~1490–1501 frontier); P(Astra appears by Dec 31 2026) ≈ 0.35–0.65 (central ~0.5); combined estimate ≈ 0.35–0.55, point estimate ~0.45–0.50.
- Multiple LLM-tracker sites disagree meaningfully on current top score (1483 vs. 1501 vs. 1525), reflecting real measurement/methodology noise on the leaderboard itself.
# Directional lean per outcome
- **Yes**: Conditional score bar (1490) is modest relative to current frontier (~1490–1525 cluster); historical debut jumps are usually positive (mean +26.6). Supports Yes *if* Astra reaches the leaderboard.
- **No**: Dominant risk is non-appearance/timing — safety pause, undecided branding, no release date, and requirement to appear (non-AutoEval) by Dec 31 2026. Polymarket's steep 30-day decline (-27.5pts) suggests market increasingly doubts timely, qualifying release. Naming ambiguity (GPT-6 vs Astra) adds resolution risk.
# Gaps / unknowns
- No primary lmarena.ai leaderboard snapshot was retrieved; current top score estimates conflict (1483–1525).
- No explicit Kalshi-direct price for this ticker was returned (only Polymarket cross-market data available).
- No sibling threshold markets found to triangulate distribution shape.
- Style-control vs. no-style-control score gap not quantified.
- No visibility into OpenAI's actual internal timeline post safety-pause.
# Calibration anchors
- Polymarket YES price (cross-market anchor): 34.5%, recent range 22.5–62%.
- Code-execution model point estimate: ~0.45–0.50.
- Historical base rate: most frontier debuts land within ±20-30 Elo of prior leader; conditional-on-appearance clearance of 1490 is very likely (~93-96%), so probability is primarily gated by appearance/timing, not score level.