Establish the database's current cumulative total and recent per-month intake, extrapolate the growth curve to late August 2026, then discount for summer court slowdown and for incomplete backfill at the resolution snapshot; the threshold sub-questions (sq2/sq3) plus the trend and lag sub-questions blend via weighted average into a median and spread for the numeric distribution.
## Cross-Market Signals ### No signal found
1. [sq1 | web_search | STRONG cred 85 | UP | VERY_RECENT] Database grew from ~200 cases mid-2025 to 719 (Jan 2026), 1,227 (early Apr 2026), 1,598 (Jun 9, 2026), 1,668 (Jul 2, 2026), and 1,809 (late Jul 2026). 2. [sq1 | web_search | STRONG cred 80 | UP | VERY_RECENT] Year-over-year comparison implies roughly a 6-9x increase in cumulative entries between mid-2025 and mid-2026, i.e., strongly positive YoY intake growth. 3. [sq1 | web_search | MODERATE cred 65 | DOWN | VERY_RECENT] Implied addition rate between Jun 9 and Jul 2, 2026 was ~3 entries/day, below the ~5.5-7.8/day pace of Jan-Jun 2026. 4. [sq1 | web_search | MODERATE cred 70 | UP | VERY_RECENT] Between Jul 2 and late July 2026 the site total rose from 1,668 to 1,809, about 5.6 new entries per day. 5. [sq1 | web_search | MODERATE cred 88 | NEUTRAL | VERY_RECENT] Question description quotes the site as saying '1730 cases identified so far,' consistent with a snapshot between early and late July 2026. 6. [sq1 | web_search | MODERATE cred 70 | UP | DATED] More than 300 US federal judges had adopted standing orders or local rules on generative AI in filings as of early 2026, increasing judicial detection and documentation. 7. [sq1 | article_search | WEAK cred 75 | NEUTRAL | DATED] Judges themselves increasingly use AI to draft rulings and prepare hearings, per April 2026 Washington Post reporting. 8. [sq2 | code_execution | MODERATE cred 40 | NEUTRAL | VERY_RECENT] Monte Carlo combining exponential growth, seasonality, and reporting lag gives median 61, mean 85, P(>=40)=0.69, P(>=100)=0.28 for the half-month count. 9. [sq2 | web_search | MODERATE cred 55 | UP | VERY_RECENT] At ~5 new entries per day overall database growth, a 16-day window would eventually accumulate roughly 80 entries if all entries were newly dated rather than backfill of older decisions. 10. [sq3 | web_search | WEAK cred 45 | DOWN | VERY_RECENT] No source reports any single half-month period in the database exceeding 100 dated cases as of mid-2026. 11. [sq4 | web_search | STRONG cred 90 | UP | VERY_RECENT] The site itself states it 'is a work in progress and will expand as new examples emerge,' confirming ongoing backfill of newly discovered decisions. 12. [sq4 | article_search | WEAK cred 60 | NEUTRAL | VERY_RECENT] Article search returned no coverage of the database's per-period intake, reporting lag, or seasonal patterns; results were largely unrelated court stories. ## Cross-Market Signals ### No signal found Information gaps: - No direct per-half-month dated-case counts from the database (e.g., first-half August 2025/2026 baselines) - No measurement of the database's typical reporting lag (days from decision date to entry) - Unknown resolution/snapshot date relative to Sept 1, 2026 - No data on August court-recess seasonality effects on this database Key uncertainties: - How much of the ~5/day intake is newly dated vs. backfilled older cases - Severity of late-August court slowdown in US filings/decisions - Whether growth is still accelerating or beginning to plateau mid-2026 - Magnitude of undercount at an early snapshot
You are an elite superforecaster. Produce a probability distribution over the answer to this Metaculus numeric question.
## Question
How many cases will Damien Charlotin's AI Hallucination Cases database record for the second half of August 2026?
## Description / Resolution Criteria
## Description
According to the resolution source: "This database tracks legal decisions1 in cases where generative AI produced hallucinated content – typically fake citations, but also other types of AI-generated arguments. It does not track the (necessarily wider) universe of all fake citations or use of AI in court filings.
While seeking to be exhaustive (1730 cases identified so far), it is a work in progress and will expand as new examples emerge. This database has been featured in news media, and indeed in several decisions dealing with hallucinated material.2"
`{"format": "bot_tournament_question", "info": {"hash_id": "51b244da24087f3e", "sheet_id": "128"}}`
## Resolution Criteria
This question resolves as the number of cases recorded by the [AI Hallucination Cases](https://www.damiencharlotin.com/hallucinations/?graphs=0&q=&sort_by=-date&period_idx=0&legal_fields=) with a date that is after August 15, 2026 and before September 1, 2026.
## Range
The answer must be a number in [9.5, 100.5] (units: cases).
## Sub-question decomposition (planner)
- (w=0.30) Will the AI Hallucination Cases database's monthly intake rate (cases dated within a given half-month) still be growing year-over-year as of mid-2026, rather than plateauing or declining? — The central driver: the database has grown exponentially through 2025; whether growth continues, plateaus, or saturates
- (w=0.25) Will the number of cases dated in the second half of August 2026 be at least 40 at resolution time? — Anchors the low end of the distribution given late-2025 half-month counts were roughly in the 30-70 range and rising.
- (w=0.25) Will the number of cases dated in the second half of August 2026 be at least 100 at resolution time? — Anchors the central/high end if the observed doubling-every-few-months trend persists into 2026.
- (w=0.20) Will resolution occur soon enough after September 1, 2026 that substantial backfill of late-August entries is still missing (i.e., recorded count materially undercounts eventual total)? — Entries are added with a lag of weeks to months; a resolution snapshot taken shortly after the period will show fewer ca
## Synthesized evidence
1. [sq1 | web_search | STRONG cred 85 | UP | VERY_RECENT] Database grew from ~200 cases mid-2025 to 719 (Jan 2026), 1,227 (early Apr 2026), 1,598 (Jun 9, 2026), 1,668 (Jul 2, 2026), and 1,809 (late Jul 2026).
2. [sq1 | web_search | STRONG cred 80 | UP | VERY_RECENT] Year-over-year comparison implies roughly a 6-9x increase in cumulative entries between mid-2025 and mid-2026, i.e., strongly positive YoY intake growth.
3. [sq1 | web_search | MODERATE cred 65 | DOWN | VERY_RECENT] Implied addition rate between Jun 9 and Jul 2, 2026 was ~3 entries/day, below the ~5.5-7.8/day pace of Jan-Jun 2026.
4. [sq1 | web_search | MODERATE cred 70 | UP | VERY_RECENT] Between Jul 2 and late July 2026 the site total rose from 1,668 to 1,809, about 5.6 new entries per day.
5. [sq1 | web_search | MODERATE cred 88 | NEUTRAL | VERY_RECENT] Question description quotes the site as saying '1730 cases identified so far,' consistent with a snapshot between early and late July 2026.
6. [sq1 | web_search | MODERATE cred 70 | UP | DATED] More than 300 US federal judges had adopted standing orders or local rules on generative AI in filings as of early 2026, increasing judicial detection and documentation.
7. [sq1 | article_search | WEAK cred 75 | NEUTRAL | DATED] Judges themselves increasingly use AI to draft rulings and prepare hearings, per April 2026 Washington Post reporting.
8. [sq2 | code_execution | MODERATE cred 40 | NEUTRAL | VERY_RECENT] Monte Carlo combining exponential growth, seasonality, and reporting lag gives median 61, mean 85, P(>=40)=0.69, P(>=100)=0.28 for the half-month count.
9. [sq2 | web_search | MODERATE cred 55 | UP | VERY_RECENT] At ~5 new entries per day overall database growth, a 16-day window would eventually accumulate roughly 80 entries if all entries were newly dated rather than backfill of older decisions.
10. [sq3 | web_search | WEAK cred 45 | DOWN | VERY_RECENT] No source reports any single half-month period in the database exceeding 100 dated cases as of mid-2026.
11. [sq4 | web_search | STRONG cred 90 | UP | VERY_RECENT] The site itself states it 'is a work in progress and will expand as new examples emerge,' confirming ongoing backfill of newly discovered decisions.
12. [sq4 | article_search | WEAK cred 60 | NEUTRAL | VERY_RECENT] Article search returned no coverage of the database's per-period intake, reporting lag, or seasonal patterns; results were largely unrelated court stories.
## Cross-Market Signals
### No signal found
Information gaps:
- No direct per-half-month dated-case counts from the database (e.g., first-half August 2025/2026 baselines)
- No measurement of the database's typical reporting lag (days from decision date to entry)
- Unknown resolution/snapshot date relative to Sept 1, 2026
- No data on August court-recess seasonality effects on this database
Key uncertainties:
- How much of the ~5/day intake is newly dated vs. backfilled older cases
- Severity of late-August court slowdown in US filings/decisions
- Whether growth is still accelerating or beginning to plateau mid-2026
- Magnitude of undercount at an early snapshot
## Required pre-forecast walkthrough
Before giving percentiles, address these explicitly in your rationale:
(a) The time left until the question resolves.
(b) The outcome if NOTHING changes from today (the status quo value).
(c) The outcome if the CURRENT TREND continues.
(d) The expectations of experts / markets / base rates.
(e) A plausible scenario that produces a LOW outcome (near p10).
(f) A plausible scenario that produces a HIGH outcome (near p90).
## Calibration guidance
- **Be humble about tails.** Good forecasters set WIDE 90/10 intervals to account for unknown unknowns. Narrow tails get punished by the log score far more than slightly-biased medians.
- **Status quo anchoring.** The p50 should be close to the status quo value unless you have strong evidence of a trend.
- Don't pile mass at one value — if you're tempted, widen the spread by 20-50%.
- **Anchor on markets/experts.** If liquid market prices, analyst forecasts, or community percentiles appear in the evidence, center your distribution on them and widen — don't override a liquid market without specific evidence it lacks.
- **Relative-return / spread questions ("how much will X's return exceed Y's").** A near-zero median is usually right, but size the TAILS to the more VOLATILE leg, not to a generic 2-3pp spread. Two broad equity indices (e.g. Nasdaq-100 vs S&P 500) do stay within roughly ±2-3pp over a two-week window. But when one leg is a commodity (crude oil, gold) or a single high-beta stock (e.g. Nvidia), the two-week realized spread regularly reaches ±10pp or more — crude-vs-equity spreads have hit ±18pp. In those cases set p10/p90 to at least ±10pp (±15pp or wider for oil) while keeping the median near the status-quo spread; the 2-3pp intuition applies only to index-vs-index.
- **Heavy-tailed one-off magnitudes (opening-weekend box office, product-launch or debut numbers).** Pre-event tracking for these routinely misses by ±40% or more, and results land both far above and far below the consensus. Do NOT treat the displayed [min, max] range as the plausible range or clamp p10/p90 to its edges: the outcome can fall outside the range in either direction (recent opening weekends resolved ~7% above a displayed ceiling in one case and roughly half the displayed floor in another). Set p10/p90 wide enough that a 2x-or-0.5x surprise versus your central estimate still lands inside the interval, and when concrete current tracking disagrees with historical/franchise base rates, weight the tracking.
## Output
Return ONLY valid JSON, no markdown fences:
{
"rationale": "<address (a) through (f) above, then state your distribution — 6-10 sentences total>",
"p10": <number>,
"p25": <number>,
"p50": <number>,
"p75": <number>,
"p90": <number>
}
Percentiles MUST be in non-decreasing order (p10 <= p25 <= p50 <= p75 <= p90).
{
"forecast": {
"evidence_brief": "1. [sq1 | web_search | STRONG cred 85 | UP | VERY_RECENT] Database grew from ~200 cases mid-2025 to 719 (Jan 2026), 1,227 (early Apr 2026), 1,598 (Jun 9, 2026), 1,668 (Jul 2, 2026), and 1,809 (late Jul 2026).\n2. [sq1 | web_search | STRONG cred 80 | UP | VERY_RECENT] Year-over-year comparison implies roughly a 6-9x increase in cumulative entries between mid-2025 and mid-2026, i.e., strongly positive YoY intake growth.\n3. [sq1 | web_search | MODERATE cred 65 | DOWN | VERY_RECENT] Implied addition rate between Jun 9 and Jul 2, 2026 was ~3 entries/day, below the ~5.5-7.8/day pace of Jan-Jun 2026.\n4. [sq1 | web_search | MODERATE cred 70 | UP | VERY_RECENT] Between Jul 2 and late July 2026 the site total rose from 1,668 to 1,809, about 5.6 new entries per day.\n5. [sq1 | web_search | MODERATE cred 88 | NEUTRAL | VERY_RECENT] Question description quotes the site as saying \u00271730 cases identified so far,\u0027 consistent with a snapshot between early and late July 2026.\n6. [sq1 | web_search | MODERATE cred 70 | UP | DATED] More than 300 US federal judges had adopted standing orders or local rules on generative AI in filings as of early 2026, increasing judicial detection and documentation.\n7. [sq1 | article_search | WEAK cred 75 | NEUTRAL | DATED] Judges themselves increasingly use AI to draft rulings and prepare hearings, per April 2026 Washington Post reporting.\n8. [sq2 | code_execution | MODERATE cred 40 | NEUTRAL | VERY_RECENT] Monte Carlo combining exponential growth, seasonality, and reporting lag gives median 61, mean 85, P(\u003e=40)=0.69, P(\u003e=100)=0.28 for the half-month count.\n9. [sq2 | web_search | MODERATE cred 55 | UP | VERY_RECENT] At ~5 new entries per day overall database growth, a 16-day window would eventually accumulate roughly 80 entries if all entries were newly dated rather than backfill of older decisions.\n10. [sq3 | web_search | WEAK cred 45 | DOWN | VERY_RECENT] No source reports any single half-month period in the database exceeding 100 dated cases as of mid-2026.\n11. [sq4 | web_search | STRONG cred 90 | UP | VERY_RECENT] The site itself states it \u0027is a work in progress and will expand as new examples emerge,\u0027 confirming ongoing backfill of newly discovered decisions.\n12. [sq4 | article_search | WEAK cred 60 | NEUTRAL | VERY_RECENT] Article search returned no coverage of the database\u0027s per-period intake, reporting lag, or seasonal patterns; results were largely unrelated court stories.\n\n## Cross-Market Signals\n\n### No signal found\n\nInformation gaps:\n - No direct per-half-month dated-case counts from the database (e.g., first-half August 2025/2026 baselines)\n - No measurement of the database\u0027s typical reporting lag (days from decision date to entry)\n - Unknown resolution/snapshot date relative to Sept 1, 2026\n - No data on August court-recess seasonality effects on this database\n\nKey uncertainties:\n - How much of the ~5/day intake is newly dated vs. backfilled older cases\n - Severity of late-August court slowdown in US filings/decisions\n - Whether growth is still accelerating or beginning to plateau mid-2026\n - Magnitude of undercount at an early snapshot",
"forecast_prompt": "You are an elite superforecaster. Produce a probability distribution over the answer to this Metaculus numeric question.\n\n## Question\nHow many cases will Damien Charlotin\u0027s AI Hallucination Cases database record for the second half of August 2026?\n\n## Description / Resolution Criteria\n## Description\nAccording to the resolution source: \"This database tracks legal decisions1 in cases where generative AI produced hallucinated content \u2013 typically fake citations, but also other types of AI-generated arguments. It does not track the (necessarily wider) universe of all fake citations or use of AI in court filings.\n\nWhile seeking to be exhaustive (1730 cases identified so far), it is a work in progress and will expand as new examples emerge. This database has been featured in news media, and indeed in several decisions dealing with hallucinated material.2\"\n\n`{\"format\": \"bot_tournament_question\", \"info\": {\"hash_id\": \"51b244da24087f3e\", \"sheet_id\": \"128\"}}`\n\n## Resolution Criteria\nThis question resolves as the number of cases recorded by the [AI Hallucination Cases](https://www.damiencharlotin.com/hallucinations/?graphs=0\u0026q=\u0026sort_by=-date\u0026period_idx=0\u0026legal_fields=) with a date that is after August 15, 2026 and before September 1, 2026.\n\n## Range\nThe answer must be a number in [9.5, 100.5] (units: cases).\n\n## Sub-question decomposition (planner)\n- (w=0.30) Will the AI Hallucination Cases database\u0027s monthly intake rate (cases dated within a given half-month) still be growing year-over-year as of mid-2026, rather than plateauing or declining? \u2014 The central driver: the database has grown exponentially through 2025; whether growth continues, plateaus, or saturates \n- (w=0.25) Will the number of cases dated in the second half of August 2026 be at least 40 at resolution time? \u2014 Anchors the low end of the distribution given late-2025 half-month counts were roughly in the 30-70 range and rising.\n- (w=0.25) Will the number of cases dated in the second half of August 2026 be at least 100 at resolution time? \u2014 Anchors the central/high end if the observed doubling-every-few-months trend persists into 2026.\n- (w=0.20) Will resolution occur soon enough after September 1, 2026 that substantial backfill of late-August entries is still missing (i.e., recorded count materially undercounts eventual total)? \u2014 Entries are added with a lag of weeks to months; a resolution snapshot taken shortly after the period will show fewer ca\n\n## Synthesized evidence\n1. [sq1 | web_search | STRONG cred 85 | UP | VERY_RECENT] Database grew from ~200 cases mid-2025 to 719 (Jan 2026), 1,227 (early Apr 2026), 1,598 (Jun 9, 2026), 1,668 (Jul 2, 2026), and 1,809 (late Jul 2026).\n2. [sq1 | web_search | STRONG cred 80 | UP | VERY_RECENT] Year-over-year comparison implies roughly a 6-9x increase in cumulative entries between mid-2025 and mid-2026, i.e., strongly positive YoY intake growth.\n3. [sq1 | web_search | MODERATE cred 65 | DOWN | VERY_RECENT] Implied addition rate between Jun 9 and Jul 2, 2026 was ~3 entries/day, below the ~5.5-7.8/day pace of Jan-Jun 2026.\n4. [sq1 | web_search | MODERATE cred 70 | UP | VERY_RECENT] Between Jul 2 and late July 2026 the site total rose from 1,668 to 1,809, about 5.6 new entries per day.\n5. [sq1 | web_search | MODERATE cred 88 | NEUTRAL | VERY_RECENT] Question description quotes the site as saying \u00271730 cases identified so far,\u0027 consistent with a snapshot between early and late July 2026.\n6. [sq1 | web_search | MODERATE cred 70 | UP | DATED] More than 300 US federal judges had adopted standing orders or local rules on generative AI in filings as of early 2026, increasing judicial detection and documentation.\n7. [sq1 | article_search | WEAK cred 75 | NEUTRAL | DATED] Judges themselves increasingly use AI to draft rulings and prepare hearings, per April 2026 Washington Post reporting.\n8. [sq2 | code_execution | MODERATE cred 40 | NEUTRAL | VERY_RECENT] Monte Carlo combining exponential growth, seasonality, and reporting lag gives median 61, mean 85, P(\u003e=40)=0.69, P(\u003e=100)=0.28 for the half-month count.\n9. [sq2 | web_search | MODERATE cred 55 | UP | VERY_RECENT] At ~5 new entries per day overall database growth, a 16-day window would eventually accumulate roughly 80 entries if all entries were newly dated rather than backfill of older decisions.\n10. [sq3 | web_search | WEAK cred 45 | DOWN | VERY_RECENT] No source reports any single half-month period in the database exceeding 100 dated cases as of mid-2026.\n11. [sq4 | web_search | STRONG cred 90 | UP | VERY_RECENT] The site itself states it \u0027is a work in progress and will expand as new examples emerge,\u0027 confirming ongoing backfill of newly discovered decisions.\n12. [sq4 | article_search | WEAK cred 60 | NEUTRAL | VERY_RECENT] Article search returned no coverage of the database\u0027s per-period intake, reporting lag, or seasonal patterns; results were largely unrelated court stories.\n\n## Cross-Market Signals\n\n### No signal found\n\nInformation gaps:\n - No direct per-half-month dated-case counts from the database (e.g., first-half August 2025/2026 baselines)\n - No measurement of the database\u0027s typical reporting lag (days from decision date to entry)\n - Unknown resolution/snapshot date relative to Sept 1, 2026\n - No data on August court-recess seasonality effects on this database\n\nKey uncertainties:\n - How much of the ~5/day intake is newly dated vs. backfilled older cases\n - Severity of late-August court slowdown in US filings/decisions\n - Whether growth is still accelerating or beginning to plateau mid-2026\n - Magnitude of undercount at an early snapshot\n\n## Required pre-forecast walkthrough\n\nBefore giving percentiles, address these explicitly in your rationale:\n (a) The time left until the question resolves.\n (b) The outcome if NOTHING changes from today (the status quo value).\n (c) The outcome if the CURRENT TREND continues.\n (d) The expectations of experts / markets / base rates.\n (e) A plausible scenario that produces a LOW outcome (near p10).\n (f) A plausible scenario that produces a HIGH outcome (near p90).\n\n## Calibration guidance\n\n- **Be humble about tails.** Good forecasters set WIDE 90/10 intervals to account for unknown unknowns. Narrow tails get punished by the log score far more than slightly-biased medians.\n- **Status quo anchoring.** The p50 should be close to the status quo value unless you have strong evidence of a trend.\n- Don\u0027t pile mass at one value \u2014 if you\u0027re tempted, widen the spread by 20-50%.\n- **Anchor on markets/experts.** If liquid market prices, analyst forecasts, or community percentiles appear in the evidence, center your distribution on them and widen \u2014 don\u0027t override a liquid market without specific evidence it lacks.\n- **Relative-return / spread questions (\"how much will X\u0027s return exceed Y\u0027s\").** A near-zero median is usually right, but size the TAILS to the more VOLATILE leg, not to a generic 2-3pp spread. Two broad equity indices (e.g. Nasdaq-100 vs S\u0026P 500) do stay within roughly \u00b12-3pp over a two-week window. But when one leg is a commodity (crude oil, gold) or a single high-beta stock (e.g. Nvidia), the two-week realized spread regularly reaches \u00b110pp or more \u2014 crude-vs-equity spreads have hit \u00b118pp. In those cases set p10/p90 to at least \u00b110pp (\u00b115pp or wider for oil) while keeping the median near the status-quo spread; the 2-3pp intuition applies only to index-vs-index.\n- **Heavy-tailed one-off magnitudes (opening-weekend box office, product-launch or debut numbers).** Pre-event tracking for these routinely misses by \u00b140% or more, and results land both far above and far below the consensus. Do NOT treat the displayed [min, max] range as the plausible range or clamp p10/p90 to its edges: the outcome can fall outside the range in either direction (recent opening weekends resolved ~7% above a displayed ceiling in one case and roughly half the displayed floor in another). Set p10/p90 wide enough that a 2x-or-0.5x surprise versus your central estimate still lands inside the interval, and when concrete current tracking disagrees with historical/franchise base rates, weight the tracking.\n\n## Output\n\nReturn ONLY valid JSON, no markdown fences:\n\n{\n \"rationale\": \"\u003caddress (a) through (f) above, then state your distribution \u2014 6-10 sentences total\u003e\",\n \"p10\": \u003cnumber\u003e,\n \"p25\": \u003cnumber\u003e,\n \"p50\": \u003cnumber\u003e,\n \"p75\": \u003cnumber\u003e,\n \"p90\": \u003cnumber\u003e\n}\n\nPercentiles MUST be in non-decreasing order (p10 \u003c= p25 \u003c= p50 \u003c= p75 \u003c= p90).\n",
"models": [
"opus",
"secondary"
],
"p50": 58.5,
"percentiles": {
"0.1": 24.5,
"0.25": 39.5,
"0.5": 58.5,
"0.75": 80.5,
"0.9": 98.0
},
"rationale": "(a) The question covers Aug 16-31, 2026, with resolution presumably at or shortly after Sept 1, 2026 \u2014 roughly a month out from the latest evidence snapshot (late July 2026). (b) Status quo: the database was adding ~5-5.6 entries/day overall in July 2026 (1,668 on Jul 2 \u2192 1,809 late July), which over a 16-day window implies ~80-90 entries added, but only a fraction of those are dated within that same window (the rest are backfill of older decisions), so a mature half-month count is plausibly ~50-80. (c) If the current trend of continued (though possibly decelerating) growth holds, the eventual Aug 16-31 dated count could be 70-100, but the count *recorded at resolution* will be materially undercounted because entries typically lag decision dates by weeks; a snapshot within days of Sept 1 could capture only ~40-60% of the eventual total. (d) The provided Monte Carlo (median 61, mean 85, P\u226540=0.69, P\u2265100=0.28) and the question\u0027s own [9.5, 100.5] range both point to a center in the 50-70 band; no market signal exists. (e) Low scenario (~p10 \u2248 22): August court recess/summer slowdown plus a snapshot taken immediately after Sept 1 with heavy reporting lag, and a plateauing intake as sanctions deterrence and AI tool citation-checking reduce new hallucination cases. (f) High scenario (~p90 \u2248 97+): continued acceleration in detection (300+ standing orders, wide media attention), a late resolution date allowing substantial backfill, and no meaningful August seasonality \u2014 pushing past 100 and clipping at the ceiling. I center near 55 with a wide, right-skewed spread reflecting deep uncertainty about lag fraction and snapshot timing."
},
"plan": {
"combination_logic": "weighted_average",
"domain": "legal",
"n_sub_qs": 4,
"n_tools": 4,
"reasoning_approach": "Establish the database\u0027s current cumulative total and recent per-month intake, extrapolate the growth curve to late August 2026, then discount for summer court slowdown and for incomplete backfill at the resolution snapshot; the threshold sub-questions (sq2/sq3) plus the trend and lag sub-questions blend via weighted average into a median and spread for the numeric distribution.",
"sub_questions": [
{
"id": "sq1",
"question": "Will the AI Hallucination Cases database\u0027s monthly intake rate (cases dated within a given half-month) still be growing year-over-year as of mid-2026, rather than plateauing or declining?",
"rationale": "The central driver: the database has grown exponentially through 2025; whether growth continues, plateaus, or saturates sets the level for Aug 2026.",
"weight": 0.3
},
{
"id": "sq2",
"question": "Will the number of cases dated in the second half of August 2026 be at least 40 at resolution time?",
"rationale": "Anchors the low end of the distribution given late-2025 half-month counts were roughly in the 30-70 range and rising.",
"weight": 0.25
},
{
"id": "sq3",
"question": "Will the number of cases dated in the second half of August 2026 be at least 100 at resolution time?",
"rationale": "Anchors the central/high end if the observed doubling-every-few-months trend persists into 2026.",
"weight": 0.25
},
{
"id": "sq4",
"question": "Will resolution occur soon enough after September 1, 2026 that substantial backfill of late-August entries is still missing (i.e., recorded count materially undercounts eventual total)?",
"rationale": "Entries are added with a lag of weeks to months; a resolution snapshot taken shortly after the period will show fewer cases than the eventual steady-state count, plus late August includes summer court recess effects.",
"weight": 0.2
}
],
"tool_requests": [
{
"parameters": {
"queries": [
"Damien Charlotin AI Hallucination Cases database number of cases",
"AI hallucination cases database total 2026 fake citations court",
"damiencharlotin.com/hallucinations cases per month growth"
]
},
"target_sub_questions": [
"sq1",
"sq2",
"sq3"
],
"tool_name": "web_search"
},
{
"parameters": {
"brief": "Find the current total number of cases in Damien Charlotin\u0027s AI Hallucination Cases database (damiencharlotin.com/hallucinations), and any reported monthly/period breakdowns of new cases added during 2025 and 2026. Also look for commentary on the rate of growth (e.g., cases per month, doubling time), the geographic mix, and how quickly new decisions are added to the database after they are issued (reporting lag).",
"max_searches": 4,
"question_title": "How many cases will Damien Charlotin\u0027s AI Hallucination Cases database record for the second half of August 2026?"
},
"target_sub_questions": [
"sq1",
"sq4"
],
"tool_name": "claude_news"
},
{
"parameters": {
"lookback_days": 120,
"queries": [
"AI hallucinated citations court sanctions database Charlotin",
"fake AI citations lawyer sanctioned 2026",
"courts AI fabricated case citations rising number"
]
},
"target_sub_questions": [
"sq1",
"sq3"
],
"tool_name": "article_search"
},
{
"parameters": {
"code": "import numpy as np\n# Model half-month intake for Aug 16-31, 2026 from exponential growth in AI hallucination case reporting.\n# Known anchor: ~1730 cumulative cases as of question writing (roughly early 2026).\n# Historical pattern (approx, from public reporting): monthly new cases ~5 (early 2025) -\u003e ~30 (mid 2025) -\u003e ~120-200 (late 2025/early 2026).\nnp.random.seed(0)\nN=200000\n# monthly rate in early 2026 (cases per month), uncertain\nbase_monthly = np.random.lognormal(np.log(150),0.45,N)\n# months from early 2026 to Aug 2026 ~ 7\nmonths=7\n# monthly growth factor, with saturation possible\ng = np.random.normal(1.06,0.10,N) # per-month multiplicative growth, can be \u003c1\nfuture_monthly = base_monthly*np.power(np.clip(g,0.75,1.35),months)\nhalf_month = future_monthly*0.5\n# late-August court recess / summer slowdown factor\nseason = np.random.normal(0.85,0.12,N)\n# backfill incompleteness at resolution snapshot\nbackfill = np.random.beta(6,3,N) # fraction captured\nres = half_month*np.clip(season,0.5,1.2)*backfill\nres = np.clip(res,0,None)\nqs=[2.5,10,25,50,75,90,97.5]\nprint({q: round(float(np.percentile(res,q)),1) for q in qs})\nprint(\u0027P(\u003e=40)\u0027, float((res\u003e=40).mean()), \u0027P(\u003e=100)\u0027, float((res\u003e=100).mean()), \u0027P(\u003e=200)\u0027, float((res\u003e=200).mean()))\nprint(\u0027mean\u0027, float(res.mean()))",
"description": "Monte Carlo of half-month case count combining exponential growth, seasonality, and reporting-lag backfill"
},
"target_sub_questions": [
"sq2",
"sq3",
"sq4"
],
"tool_name": "code_execution"
}
]
},
"question": {
"close_time": "2026-07-27T21:00:00Z",
"description": "## Description\nAccording to the resolution source: \"This database tracks legal decisions1 in cases where generative AI produced hallucinated content \u2013 typically fake citations, but also other types of AI-generated arguments. It does not track the (necessarily wider) universe of all fake citations or use of AI in court filings.\n\nWhile seeking to be exhaustive (1730 cases identified so far), it is a work in progress and will expand as new examples emerge. This database has been featured in news media, and indeed in several decisions dealing with hallucinated material.2\"\n\n`{\"format\": \"bot_tournament_question\", \"info\": {\"hash_id\": \"51b244da24087f3e\", \"sheet_id\": \"128\"}}`\n\n## Resolution Criteria\nThis question resolves as the number of cases recorded by the [AI Hallucination Cases](https://www.damiencharlotin.com/hallucinations/?graphs=0\u0026q=\u0026sort_by=-date\u0026period_idx=0\u0026legal_fields=) with a date that is after August 15, 2026 and before September 1, 2026.",
"open_lower": true,
"open_upper": true,
"q_max": 100.5,
"q_min": 9.5,
"question_type": "numeric",
"title": "How many cases will Damien Charlotin\u0027s AI Hallucination Cases database record for the second half of August 2026?",
"units": "cases"
},
"research": {
"cross_market_brief": "## Cross-Market Signals\n\n### No signal found",
"errors": [],
"has_cross_market": true,
"n_errors": 0,
"n_tools": 4,
"tools": [
{
"elapsed_s": 31.59,
"error": null,
"success": true,
"summary": "0 results",
"tool_name": "web_search"
},
{
"elapsed_s": 36.43,
"error": null,
"success": true,
"summary": "11 citations",
"tool_name": "claude_news"
},
{
"elapsed_s": 0.84,
"error": null,
"success": true,
"summary": "18 articles",
"tool_name": "article_search"
},
{
"elapsed_s": 0.29,
"error": null,
"success": true,
"summary": "Monte Carlo of half-month case count combining exponential g",
"tool_name": "code_execution"
}
]
},
"synthesis": {
"evidence": [
{
"claim": "Database grew from ~200 cases mid-2025 to 719 (Jan 2026), 1,227 (early Apr 2026), 1,598 (Jun 9, 2026), 1,668 (Jul 2, 2026), and 1,809 (late Jul 2026).",
"credibility": 85,
"direction": "UP",
"priced_in": true,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "STRONG",
"sub_question_id": "sq1"
},
{
"claim": "Year-over-year comparison implies roughly a 6-9x increase in cumulative entries between mid-2025 and mid-2026, i.e., strongly positive YoY intake growth.",
"credibility": 80,
"direction": "UP",
"priced_in": true,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "STRONG",
"sub_question_id": "sq1"
},
{
"claim": "Implied addition rate between Jun 9 and Jul 2, 2026 was ~3 entries/day, below the ~5.5-7.8/day pace of Jan-Jun 2026.",
"credibility": 65,
"direction": "DOWN",
"priced_in": false,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "MODERATE",
"sub_question_id": "sq1"
},
{
"claim": "Between Jul 2 and late July 2026 the site total rose from 1,668 to 1,809, about 5.6 new entries per day.",
"credibility": 70,
"direction": "UP",
"priced_in": false,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "MODERATE",
"sub_question_id": "sq1"
},
{
"claim": "Question description quotes the site as saying \u00271730 cases identified so far,\u0027 consistent with a snapshot between early and late July 2026.",
"credibility": 88,
"direction": "NEUTRAL",
"priced_in": true,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "MODERATE",
"sub_question_id": "sq1"
},
{
"claim": "More than 300 US federal judges had adopted standing orders or local rules on generative AI in filings as of early 2026, increasing judicial detection and documentation.",
"credibility": 70,
"direction": "UP",
"priced_in": true,
"recency": "DATED",
"source": "web_search",
"strength": "MODERATE",
"sub_question_id": "sq1"
},
{
"claim": "Judges themselves increasingly use AI to draft rulings and prepare hearings, per April 2026 Washington Post reporting.",
"credibility": 75,
"direction": "NEUTRAL",
"priced_in": true,
"recency": "DATED",
"source": "article_search",
"strength": "WEAK",
"sub_question_id": "sq1"
},
{
"claim": "Monte Carlo combining exponential growth, seasonality, and reporting lag gives median 61, mean 85, P(\u003e=40)=0.69, P(\u003e=100)=0.28 for the half-month count.",
"credibility": 40,
"direction": "NEUTRAL",
"priced_in": false,
"recency": "VERY_RECENT",
"source": "code_execution",
"strength": "MODERATE",
"sub_question_id": "sq2"
},
{
"claim": "At ~5 new entries per day overall database growth, a 16-day window would eventually accumulate roughly 80 entries if all entries were newly dated rather than backfill of older decisions.",
"credibility": 55,
"direction": "UP",
"priced_in": false,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "MODERATE",
"sub_question_id": "sq2"
},
{
"claim": "No source reports any single half-month period in the database exceeding 100 dated cases as of mid-2026.",
"credibility": 45,
"direction": "DOWN",
"priced_in": false,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "WEAK",
"sub_question_id": "sq3"
},
{
"claim": "The site itself states it \u0027is a work in progress and will expand as new examples emerge,\u0027 confirming ongoing backfill of newly discovered decisions.",
"credibility": 90,
"direction": "UP",
"priced_in": true,
"recency": "VERY_RECENT",
"source": "web_search",
"strength": "STRONG",
"sub_question_id": "sq4"
},
{
"claim": "Article search returned no coverage of the database\u0027s per-period intake, reporting lag, or seasonal patterns; results were largely unrelated court stories.",
"credibility": 60,
"direction": "NEUTRAL",
"priced_in": true,
"recency": "VERY_RECENT",
"source": "article_search",
"strength": "WEAK",
"sub_question_id": "sq4"
}
],
"information_gaps": [
"No direct per-half-month dated-case counts from the database (e.g., first-half August 2025/2026 baselines)",
"No measurement of the database\u0027s typical reporting lag (days from decision date to entry)",
"Unknown resolution/snapshot date relative to Sept 1, 2026",
"No data on August court-recess seasonality effects on this database"
],
"key_uncertainties": [
"How much of the ~5/day intake is newly dated vs. backfilled older cases",
"Severity of late-August court slowdown in US filings/decisions",
"Whether growth is still accelerating or beginning to plateau mid-2026",
"Magnitude of undercount at an early snapshot"
],
"n_evidence": 12
},
"timings": {
"forecast": 54.26,
"plan": 28.29,
"research": 36.44,
"synthesis": 28.9
}
}