Prediction Markets

How AI-Written Betting Markets Fail — the Authoring Doctrine (2026)

AI can write thousands of market questions, and most failures are structural and predictable. Seven failure modes, the doctrine that prevents them, and a checklist.

AI-Written MarketsPrediction MarketsMarket AuthoringMarket ResolutionAdkuu Pulse
How AI-Written Betting Markets Fail — the Authoring Doctrine (2026)

TL;DR

A language model can draft thousands of event-contract questions a day. Meta's planned prediction app would use Llama "to automatically generate questions from trending topics" and resolve markets in "near real-time", according to internal documents NPR reported on 24 June 2026. The failures are not random. In a study first published on 30 January 2026, web-research agents generated 14,073 candidate questions; 1,793 passed the verifiers, deduplication removed 294 more, and human experts still hard-rejected 8.1% of a reviewed sample as "ambiguous or otherwise flawed". The seven failures below each have a structural cause and a rule that prevents them. The fix is an authoring doctrine, not a better prompt.

Key takeaways

  • The failure modes are structural. They come from how an event unfolds in time, how evidence appears, and how a title compresses a rule, not from model quality.
  • The dated record shows each one: a Polymarket market on a deal "before April 2025" resolved YES in March 2025 while the agreement was signed on 30 April; a Kalshi title that promised an event while its rule settled on a price; a market settled on an evidence rule that appeared on its page only after the deadline.
  • Human-curated platforms are not immune. The January 2026 study puts Metaculus's historical annulment rate at about 8%, against about 3.9% for its automated pipeline.
  • The doctrine has five parts: event model first; deadlines after the whole timeline; source pinned before publishing; title and rule in lockstep; every generated market gated.
  • Adkuu Pulse's "Magic Questions" turn an operator's free-text idea into a fully structured question with resolution criteria, settled by the same multi-tier resolution as the rest of the feed.

Why this is a 2026 problem

On 24 June 2026 NPR's Bobby Allyn reported, from internal documents, that Meta is building a standalone prediction app, codenamed "Antwerp" and "FBForecast", in which users "will receive 'a daily virtual allotment' of 'play money' that can be used to bet on 'the outcome of future events.'" The app "will use Llama, the company's large language model, to automatically generate questions from trending topics", and "that same artificial intelligence" will resolve markets, a process the documents say "will happen in 'near real-time'". The documents describe a rebuild of Forecast, the crowdsourced prediction app Meta released in 2020 and wound down two years later, citing "the operational cost of manual question curation". Meta declined to comment; The New York Times first reported the plans. The design intent is clear: generation and resolution by the model, at the cadence of a news feed.

The best public measurement of what that produces is "Automating Forecasting Question Generation and Resolution for AI Evaluation" by Nikos Bosse and colleagues (arXiv, 30 January 2026, revised 9 March 2026). From 2,500 seeds (500 company revenue-forecast rationales, 200 Media Cloud articles, 1,800 GDELT articles) a web-research agent wrote 14,073 proto-questions, a second agent added resolution criteria, and four verifier agents scored quality, ambiguity, resolvability and difficulty. Only 1,793 questions (12.7%) passed; deduplication removed 294 of those, leaving 1,499. Experts then reviewed 149: 75.2% were accepted, 16.8% soft-rejected as "trivial or otherwise uninteresting", and 8.1% hard-rejected as "ambiguous or otherwise flawed". On resolution, three Gemini 3 Pro agents disagreed on 131 of 1,499 questions (8.7%), a human audit of 100 resolutions found 4 errors, "a 4.9% error rate", and the authors estimate an annulment rate of about 3.9%, against a historical rate they put at about 8% for Metaculus.

Adkuu's reading: the pipeline is the product. The generator produced the questions; the verifiers rejected about 87% of them, and a residue of ambiguity survived even that.

The seven failure modes

#FailureExample from the recordWhy it happensRule that prevents it
1Deadline before the timelineA Polymarket market on whether Ukraine would agree to a minerals deal "before April 2025" resolved YES in March 2025 after two dispute rounds; the agreement was signed on 30 April (Cointelegraph, 4 June 2026). The study notes that questions on institutional timelines "almost invariably resolve no"The date is taken from the headline, not from the steps the event needs: negotiation, signature, ratification, publicationBuild the event model first; set the deadline after the last step; write absolute dates with a time zone
2Premise already falseThe study's difficulty check was designed to flag forecasts below 2% or above 98% as trivial (the authors kept them for analysis); 16.8% of expert-reviewed questions were soft-rejected as "trivial or otherwise uninteresting"The generator works from a seed written before the latest newsCheck the current state at authoring; list near-certain markets only deliberately
3Unresolvable sourceA question on a North Korean missile launch required confirmation from the US Department of Defense, which never came, so it resolved NO although the launch occurred (EA Forum, 20 July 2021); the study found questions "where the admissible resolution sources did not publish the required data"The source was chosen for authority, not checked for publicationPin a source known to publish, name a fallback, and answer Metaculus's question "What if the data source you specified stops being published?"
4Two questions in oneThe Strategy market asked whether a bitcoin sale would happen by 31 May 2026 and was settled on whether confirmation arrived by then: "Confirmation achieved outside of the market's time frame does not qualify." (The Block, 4 June 2026)The event and its evidence are separate processes with separate clocksOne event, one condition; if confirmation timing matters, write it in the rule and the title
5Title drifts from the rule"Ali Khamenei out as Supreme Leader?" over a rule that settled at the last traded price before death (InGame, 4 March 2026). Polymarket's own rule: "The market title describes the market, but the rules define how it should be resolved."Titles are compressed for a feed; the rule is written separatelyGenerate the title from the rule; apply Metaculus's check that the headline avoids "making a stronger/weaker claim than the resolution criteria"
6Duplicates under different wordingThe study removed 294 near-duplicates from 1,793 verified questions, clustering embeddings at a cosine similarity of 0.85 and confirming with a modelGenerators sample the same news many timesEmbedding deduplication plus a model check before listing; one rule per event
7Resolves on a technicalityA Polymarket market titled "What will Trump say this week", listing "Ayatollah / Khamenei" for the week of 8 March 2026, ended in a $170,000 payout claim over whether the word was said (Crypto Briefing, 8 August 2026). The study warns that "the absence of Google results suggesting that this occurred does not imply that it did not occur"The evidence standard was never writtenWrite what counts as proof and the default when proof is absent; Metaculus's checklist asks whether resolution avoids "dependence on any individual or group saying particular words or phrases"

Two patterns run through the table. Failures 1, 2 and 4 are about time: the author did not model the sequence of the event, so the deadline, the premise or the evidence window was wrong. Failures 3, 5 and 7 are about the rule: the source, the title or the evidence standard was left implicit, and someone else decided it later. Failure 6 is about volume, and only matters once a generator is running.

The authoring doctrine

1. Event model first. Before writing a question, write the sequence: what has to happen, in which order, who announces it, where, and how long each step takes. Metaculus's question-writing guide says "Be concrete", "Define your terms", "Consider and account for edge-cases" and "Try to account for unknown unknowns". The event model is where those instructions become concrete: an announcement is not a signature, and a filing is not the act it describes.

2. Deadlines after the whole timeline. Set the deadline after the last step in the event model, not after the first. The guide requires that "The date or time period the question is asking about must always be explicitly mentioned in the text", written in the form "January 1, 2040" with a time zone or UTC offset, and that "The close date must be at least one hour prior to the resolution date". A deadline that falls inside the timeline produces a market that resolves NO on a technicality, or YES on a dispute.

3. Pin the source before publishing. "Use authoritative sources, when possible", and "be sure that the sources will be available at the time of resolution", with alternative sources specified otherwise. Metaculus's approval checklist asks whether "fallback sources" are listed. The study's authors allow one exception, and state its condition: "Sometimes we don't know the resolution source in advance. In many cases that's fine, as long as we can be absolutely certain that it will exist."

4. Title and rule in lockstep. The rule is the contract and the title is its summary; the two must make the same claim. The cost of getting this wrong after listing is visible in Polymarket's documentation, where a clarification clears the order book and cancels every resting order. A second reader, human or model, compares title and rule before publication and rejects any pair where one promises more.

5. Gate every market. Nothing generated reaches the feed without passing explicit checks. The study's verifier design is the published template; the table below adds the gates that follow from points 1 to 4.

GateWhat it checksPass condition
QualityWhether more research effort should improve a forecast"great" on the study's four-point scale
AmbiguityWhether two readers would resolve the question the same way"great" on the same scale
ResolvabilityWhether an agent could resolve it from available evidence"very certainly yes"
TrivialityThe current probabilityFlag below 2% or above 98%
TimelineDeadline against the event modelAfter the last step (doctrine)
LockstepTitle against ruleNo stronger or weaker claim either way (doctrine)
DeduplicationSimilarity to existing marketsEmbedding cluster at 0.85 plus a model check
SourceThe named source and fallbackKnown publisher with a schedule, or a certain-to-exist source (doctrine)

Where Adkuu Pulse fits

Adkuu Pulse is built on the same steps in two places. Its "Magic Questions" let an operator submit a free-text idea and receive a fully structured question with resolution criteria, so that the structuring happens at authoring time rather than in a dispute afterwards. Its automated multi-tier resolution (a data API first, then web search, then consensus, with human review as the fallback) then settles the market against the rule that was written. The answer pages on what an AI-written betting market is and how automated market resolution works describe both; the product page is Adkuu Pulse. The companion posts cover how prediction markets settle and why disputes happen and how to price prediction markets as fixed odds.

The authoring checklist

  1. Is the event model written: the steps, their order, and who publishes each?
  2. Is the deadline after the last step, in absolute form with a time zone?
  3. Is the premise true today, and is the current probability away from 0 and 1 unless the market is deliberately trivial?
  4. Is the primary source named, known to publish and on a schedule, with a fallback in order of precedence?
  5. Is there one event and one condition, and if confirmation timing matters, is it in both the rule and the title?
  6. Does the title make exactly the claim the rule makes?
  7. Is the evidence standard written: what counts as proof, where, and the default when proof is absent?
  8. Has the market been checked against existing markets for duplicates?
  9. Is the outcome for postponement, cancellation and partial occurrence written?
  10. Has a second reader, human or model, passed the market through every gate?

FAQ

What is an AI-written betting market? A market whose question, resolution rule and often deadline were drafted by a language model from a news item or an operator's idea, rather than by a human trader. Meta's documents describe Llama generating questions from trending topics; Adkuu Pulse's "Magic Questions" structure an operator's free-text idea into a question with resolution criteria.

How often do generated questions fail? In the January 2026 study, 1,793 of 14,073 generated questions (12.7%) passed the verifiers, deduplication removed 294 more, and experts still hard-rejected 8.1% of a reviewed sample as ambiguous or flawed. Resolution of the surviving questions carried a 4.9% error rate in a human audit.

What is the single most common structural failure? In the documented cases, timing: a deadline set before the event's own timeline has played out, or an evidence window that was never written. The Ukraine minerals market and the Strategy bitcoin market are both timing failures.

How does Adkuu Pulse prevent these failures? By structuring markets, including operator-submitted "Magic Questions", into questions with resolution criteria at authoring time, and by settling it with automated multi-tier resolution (data API, then web search, then consensus) with human review as the fallback.

Sources

Last verified: October 10, 2026.