Who Settles the Bet? Prediction Markets’ Oracle Problem in 2026
AI agents place most of the bets on Polymarket, but every payout hinges on a separate layer that decides what actually happened. Inside the oracle problem still settling wins as losses in 2026.
A prediction-market trader can read the world correctly and still lose every dollar. In late May 2026, Strategy, the corporate treasury company that used to be called MicroStrategy, sold 32 Bitcoin. The sale happened, and it happened inside the window that a Polymarket contract was asking about. On the plain facts, the people who bet Yes on whether the company would sell any Bitcoin had won. The market paid No, because the filing that proved the sale arrived one day after the contract expired. Two traders are now suing Polymarket and its chief executive for at least $797,198, and their complaint is not really about Bitcoin. It is about who gets to decide what happened.
That question, the distance between an event and a market’s ruling on the event, is the oracle problem, and it is the most important unsolved piece of the prediction-market boom. Autonomous AI agents now place a large share of the orders on Polymarket, and institutional money has arrived in size, with Intercontinental Exchange, the owner of the New York Stock Exchange, committing close to $2 billion. Yet every one of those trades settles through a separate layer that decides the truth, and that layer runs on token votes, human committees, and, increasingly, artificial intelligence. This is a piece about who settles the bet, why settlement keeps going wrong, and why the next automation frontier in prediction markets is not the betting but the judging.
The trade you win and the payout you lose
Every prediction-market position carries two independent risks. The first is the risk that you read the event wrong. The second is the risk that the market resolves the event wrong. Traders and their models obsess over the first and barely price the second, yet it is the second that has produced the lawsuits, the governance fights, and the refunds. Resolution risk is uncorrelated with skill: you can be completely right about the world and still be handed a zero, because a definition was loose, a source was late, or a vote went sideways.
The reason this matters so much now is scale. Combined monthly volume across the two dominant venues, Polymarket and Kalshi, runs into the tens of billions of dollars, and analysts have projected the sector growing into a trillion-dollar market by the end of the decade. That is an enormous amount of money riding on the last step of the process, the moment a messy real-world event is converted into a clean payout. Two designs dominate that step in 2026, and both have failed in their own way: a decentralized oracle, used by Polymarket, that lets outsiders propose and vote on outcomes, and a centralized committee, used by Kalshi, that keeps the decision inside the company. Understanding the difference is the difference between a tradable market and an untradeable one.
What a prediction market actually sells
A prediction market sells a claim on a future fact. A Yes share pays $1 if the stated event occurs and $0 if it does not; the price in between is the crowd’s implied probability. The design is elegant only if one thing holds at the end: that something reliably converts the world into a binary $1 or $0. That converter is the oracle in decentralized markets and the settlement desk in regulated ones. Everything else, the order book, the AI traders, the liquidity, sits downstream of it.
Which means a contract is only ever as trustworthy as two things fixed before trading opens: its resolution source and its resolution criteria. A well-built market names an objective source (a company’s regulatory filing, an official statistic, a wire-service call) and a resolution date up front, so there is nothing left to argue about later. The trouble is that reality does not always cooperate. As Hart Lambur, the former Goldman Sachs trader who co-founded UMA, the oracle behind Polymarket, has pointed out, some markets are subjective from birth, and others become subjective only when an edge case appears between the moment a market is created and the moment it must resolve. That gap between the clean question and the messy answer is where the entire fight lives.
Inside UMA’s optimistic oracle
Polymarket does not resolve its own markets. It outsources truth to UMA’s optimistic oracle, and the word optimistic is the whole idea: someone proposes an answer, and the system assumes that answer is correct unless somebody challenges it. In practice, a proposer posts an outcome along with a bond of around $750; a two-hour challenge window opens; anyone who disagrees can dispute by posting a matching bond; and if nobody challenges, the market finalizes and winning shares redeem for $1. If it is challenged, the case escalates to UMA’s Data Verification Mechanism, a vote of UMA token holders that runs roughly 48 hours, and the loser forfeits their bond to the winner. The economic wager is that honesty is the most profitable answer, because the majority of token holders are assumed to converge on the observable truth.
After a wave of bad or premature proposals, UMA tightened the top of the funnel. In late 2025 it rolled out a managed version of the oracle that limits who may propose to a whitelist, 37 addresses made up of staff and users with strong historical accuracy, while keeping disputes open to anyone. The effect on proposal quality was real: whitelisted proposers have resolved at about 99.7% accuracy, against 85.8% for the older open pool. But note the crucial limit, because it sets up everything that follows. The reform restricted proposing, not voting. The part of the system that decides genuinely contested cases, the token-weighted vote, is unchanged.
| Stage | What happens | Timing | Stake |
|---|---|---|---|
| Proposal | A whitelisted proposer posts an outcome to the optimistic oracle | Any time after the market’s end date | About $750 bond |
| Challenge window | Anyone can dispute the proposed outcome by matching the bond | Two-hour liveness period | About $750 counter-bond |
| Undisputed settlement | With no challenge, the outcome finalizes and winning shares redeem for $1 | About two hours total | Proposer recovers bond plus a reward |
| Dispute to the DVM | A challenge escalates to a vote of UMA token holders, weighted by staked tokens | Roughly 48 hours of voting | The losing side forfeits its bond |
| Final resolution | The token vote sets the outcome; the winner also takes half the loser’s bond | Four to six days end to end | Winner recovers bond plus half the other bond |
The Strategy Bitcoin market: a $797,000 lesson in timing
Return to the case that opened this piece, because it is the cleanest illustration of resolution risk anyone has produced. The market asked whether Strategy would sell any Bitcoin by May 31, 2026, and its designated source was the company’s filings with the Securities and Exchange Commission. Strategy did sell, 32 Bitcoin worth roughly $2.5 million, between May 26 and 31. The confirming Form 8-K, however, posted on June 1. At the instant the contract expired, the sale was real but not yet public in the source the market had named. A proposer put up No, the outcome was disputed, and when it went to a UMA vote, the result came back overwhelmingly No.
William Wood and Thomas Bush then sued in New York state court, naming Polymarket, chief executive Shayne Coplan, and other executives and entities. They are seeking at least $797,198 and, tellingly, a court order barring the platform from changing settlement rules after an outcome is known. Whatever a judge decides, the underlying lesson is stark. The traders were right about the world; they lost on a definitional question the contract never nailed down, namely whether selling means the act of selling or the public confirmation of the act. No AI model, however sophisticated, could have priced that, because the ambiguity lived in the wording, not the data. This is the tail that makes serious traders nervous, and it is not an edge case; it is the recurring shape of every resolution dispute.
When the house overrules the oracle
Even a decentralized oracle turns out to have a human escape hatch, and Polymarket has used it. In one widely discussed case, UMA voters resolved a market on whether Barron Trump was involved with the DJT token as No; Polymarket then publicly overrode the result, said he was involved in some way, and refunded Yes holders, without laying out its evidence. Lambur characterized the underlying question as genuinely ambiguous, with voters settling on No as the least bad answer to a prompt that had no clean one. More than a million dollars had traded on the market.
The episode exposed a governance paradox that never really goes away. A system marketed as trustless still keeps a discretionary override, and the moment a platform can overrule its own oracle, the oracle becomes advisory rather than final. That is the same tension crypto keeps rediscovering wherever a supposedly autonomous system quietly keeps a manual switch, the debate that runs through every argument about pause buttons and off-switches in DeFi. For a trader, and especially for a machine trader, a discretionary override is unmodelable. You cannot assign a probability to a human changing their mind.
Token-weighted truth: the conflict at the center of the vote
The deeper problem with the decentralized model is who votes, and why. Disputes are settled by UMA token holders, with votes weighted by the number of tokens staked, which means the arbiter of truth is, mechanically, whoever holds the most tokens. And tokens can be bought, borrowed, and concentrated. In one episode in early 2025, a single holder amassed several million UMA across a handful of wallets and cast a large share of the vote on a multimillion-dollar market, an outcome that disputing traders said the concentration had swung. Worse, reporting through 2026 found that voters frequently hold positions in the very markets they are asked to judge. The referee is often also a player.
UMA can fairly point out that fewer than 2% of proposals are ever disputed, so the machine runs smoothly the overwhelming majority of the time. But the disputed sliver is exactly where the stakes, the ambiguity, and the motive concentrate, and the managed-oracle reform did nothing for it, because it restricted proposing rather than the vote that decides disputes. There is an economic irony underneath all of this too: UMA, the token that secures the vote, is a mid-cap asset worth only tens of millions of dollars, yet it is routinely asked to settle contracts worth far more than the token that secures them.
Kalshi’s answer: keep resolution in-house
The regulated alternative solves the conflicted-crowd problem by removing the crowd entirely. On Kalshi, a designated contract market registered with the Commodity Futures Trading Commission, every market names an official source and a resolution date before trading opens. When the event resolves, Kalshi verifies the result against that source and settles, usually within about three hours, at $1 per winning contract with no settlement fee. Disputes do not go to a token vote; they go to Kalshi’s own Outcome Review Committee under Rule 7.1, which can open a review at the company’s discretion and must reach a final, binding determination within 24 hours.
Clean and fast, and with its own conflict baked in. The Massachusetts Attorney General, Andrea Campbell, whose office became the first in the country to sue a prediction market, argued in her complaint that Kalshi writes the rules for the contract, determines the basis for settlement, and runs a process that functions entirely within its own corporate structure and does not serve as an independent intermediary. Put plainly, the venue writes the question, picks the source, judges the outcome, and pays the winner. Decentralized resolution risks a conflicted crowd; centralized resolution risks a conflicted house. Neither is obviously safer, and both are one contested market away from a headline.
Two models, two failure modes
The choice between the two designs is really crypto’s oldest argument, decentralization versus a single accountable operator, replayed on the settlement layer instead of the ledger. One spreads the decision across a crowd that can be captured; the other concentrates it in a company that can be conflicted. The table below lays the trade-offs side by side.
| Dimension | Polymarket (UMA optimistic oracle) | Kalshi (in-house committee) |
|---|---|---|
| Who proposes the result | Whitelisted proposers; anyone can dispute | Kalshi staff, against a named data source |
| Who decides a dispute | UMA token holders, votes weighted by staked tokens | Kalshi’s Outcome Review Committee (Rule 7.1) |
| Speed | About two hours undisputed, days if escalated | Most markets about three hours, up to 24 hours if reviewed |
| Transparency | Fully on-chain; every bond and vote is visible | Internal; determinations disclosed as final |
| Main conflict risk | Voters can hold positions in the market they judge | Venue writes rules, sets settlement, and pays out |
| Regulatory footing | Event contracts face the CFTC; the UMA token is the only SEC-relevant piece | CFTC-regulated designated contract market |
Why AI agents care about resolution more than anyone
Bring this back to the machines, because the machines are the reason resolution now matters at industrial scale. Autonomous agents dominate Polymarket trading: more than 30% of active wallets are AI-run, and profitable-agent rates run well ahead of humans. David Minarsch, chief executive of Valory, which builds the Polystrat agent, describes it as an autonomous AI agent that trades on Polymarket around the clock, and says agents tend to do better than humans, while cautioning that naive off-the-shelf models produce results no better than a coin-flip. These agents run through smart-contract wallets with scoped, session-based permissions, the kind of delegated wallet control that lets software sign thousands of trades without a human touching each one.
Here is the catch that no amount of model quality fixes. An agent can price probability brilliantly and still have no way to price resolution risk. It can read an SEC filing, but it cannot foresee that a market will demand public confirmation by a deadline, that a whale will swing a dispute, or that a platform will override its own oracle. Those are governance events, not data events, and they do not appear in any training set. Worse, the agent’s real edge lives in the long tail of niche, low-liquidity markets, which is exactly where the wording is loosest and the resolution is most contestable. Resolution risk is the one tail an agent cannot hedge, which is why the most disciplined desks now sort markets into two buckets: clean, objectively sourced contracts they will trade, and ambiguous ones they will not touch at any price.
Resolution goes autonomous: UMA’s Optimistic Truth Bot
The same automation wave that took over the trading is now coming for the judging. UMA has built the Optimistic Truth Bot, an AI system that fills the proposer role: it reads a market, gathers evidence, and proposes a resolution, while humans keep the dispute and voting backstop. The architecture is a router that dispatches requests to specialized solvers, one for web search, another that runs code to query APIs, and an overseer that checks the result and can restart the whole process if it looks wrong.
Across more than 3,000 markets, the bot proposed correctly about 78% of the time. That headline number hides a sharp split by difficulty: it hit 99.3% on simple sports and asset-price markets and 95% on clean yes-or-no questions, but fell to 72% on messy mention-counting markets. Only about 0.4% of its proposals drew a genuine human challenge, which UMA reads as a high bar cleared. The limitations are stated honestly and they are instructive: the bot still only posts recommendations to social media rather than committing them on-chain, and it struggles most on open-ended and geopolitical questions, the very ones humans argue about. The pattern mirrors trading exactly. AI is superb where the target is objective, and unreliable where it is subjective.
| Approach | Reported accuracy | Where it works and breaks |
|---|---|---|
| UMA Optimistic Truth Bot (AI proposer, 3,000+ markets) | About 78% overall; 99.3% on simple sports and price markets; 72% on mention-counting | Strong on objective data, weak on open-ended and geopolitical questions |
| Best single frontier model (academic study, 1,189 questions) | 82.42% | Fails on ambiguity and timing gaps |
| Confidence-weighted model ensemble | 83.43% | Barely beats one model; errors correlate across models |
| Deliberative multi-model debate | About 76% | Underperforms a single model |
| Hybrid: auto-resolve only unanimous, high-confidence cases | 97.87% on about 47% of questions | Humans still required for the contested tail |
Name the judge before the bet: the case for AI referees
If AI is going to do the judging, the interesting design question is when it should be named. Andrew Hall, the Davies Family Professor of Political Economy at Stanford’s Graduate School of Business and a senior fellow at the Hoover Institution, has laid out the most cited proposal, published through a16z crypto. Instead of resolving a market after the fact with whatever oracle is handy, he argues the market should commit to its judge in advance. In his words, at contract creation the market maker specifies not just the resolution criteria in natural language, but the exact large language model, identified by a timestamped model version, and the exact prompt that will be used to determine the outcome. The payoff, he writes, is that the entire resolution mechanism is visible and auditable before anyone places a bet: no rule changes mid-flight, no discretionary judgment calls, no backroom negotiations.
The proposal aims straight at the failures above. The Strategy and Barron disputes were, at bottom, rule changes after the fact, and pinning the model and the prompt at creation turns resolution from a discretionary act into a committed computation. The catch is twofold. First, garbage prompt in, garbage ruling out; the ambiguity does not disappear, it just moves upstream into how carefully the prompt is written. Second, models get deprecated, so a market that names a specific model today may outlive the endpoint that runs it, which forces awkward questions about version pinning and reproducibility. Even so, the direction, commit the judge before the money moves, is the strongest single idea on the table, and it is the one that most directly closes the loophole the lawsuits are built on.
Can a committee of models settle real money?
Naming one model raises an obvious follow-up: is one model enough? A 2026 study by Tarun Kota, titled Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution, tested exactly this on 1,189 real prediction-market questions. A single best-in-class model resolved 82.42% correctly; a confidence-weighted ensemble of several models did only marginally better, at 83.43%; and, counterintuitively, letting the models debate each other toward a consensus did worse, around 76%, below any single model on its own. The reason is that model errors are correlated (the paper measures the correlation between roughly 0.53 and 0.69), so stacking models that trip on the same hard questions buys very little.
The genuinely useful result is the hybrid. When the system auto-resolved only the questions where the models were unanimous and highly confident, it covered about 47% of the set at 97.87% accuracy, and escalated the rest to human review. That is the shape the entire industry is quietly converging on, and it maps cleanly onto UMA’s own Truth Bot experience: let AI clear the clean majority instantly and almost for free, and reserve scarce, expensive human judgment for the contested tail, where, as the Strategy market showed, the money and the ambiguity both live. Resolution is not going fully autonomous. It is going mostly autonomous, with a human backstop that has to be independent to be worth anything.
The regulators are asking the same question
Who resolves the bet is also, at bottom, the legal fight. In the United States, prediction-market event contracts fall to the Commodity Futures Trading Commission, not the SEC; the SEC’s reach here extends only to the tokens bolted onto the machinery, such as UMA and OLAS, not the contracts themselves. That distinction matters because the courts are now split on what these contracts even are. In August 2026 the Ninth Circuit ruled that Kalshi’s sports contracts are sports bets, not swaps, handing states authority over them, while the Third Circuit had earlier sided with Kalshi; New Jersey has since petitioned the Supreme Court to settle the split.
Underneath the jurisdictional label sits a resolution question. A sportsbook sets the odds and is your counterparty; an exchange matches traders and settles a binary outcome from a named source. Which one a given contract is depends largely on how it resolves, so the classification fight is a resolution fight wearing a legal costume. Meanwhile the federal rulebook remains unfinished; the CFTC’s event-contract rulemaking is still pending, part of a broader regulatory calendar that reset after the CLARITY Act stalled in the Senate. The September FOMC decision, when the Fed raised rates by a quarter point, was a reminder that settlement clarity matters even for macro markets: the two big venues priced the same live outcome very differently right up until it resolved.
What good resolution would look like
If the trading layer is largely solved and the resolution layer is not, the fixes are at least coming into focus. Four principles keep recurring among the people trying to fix it. First, objective named sources and resolution criteria pinned down before trading opens, so a word like sell or a phrase like wore a suit cannot be redefined after the fact. Second, AI proposers for the clean majority, with the model and prompt published in advance, along the lines of Hall’s commit-before-the-bet principle. Third, a genuinely conflict-free dispute layer, because token-weighted voting by market participants and in-house committees both fail the independence test, just in opposite directions. Fourth, appeals that carry real skin in the game and a hard ban on retroactive rule changes, which is precisely the remedy the Strategy plaintiffs are asking a court to impose.
The stakes only rise from here. With Intercontinental Exchange’s roughly $2 billion bet on Polymarket as data infrastructure, and serious projections of a trillion-dollar sector this decade, the money now depends on settlement being boringly reliable, in the same way that equities and options became investable at scale only once clearing and settlement were solved, long before anyone trusted the wrapper products built on top. Bots already bet, and machines are learning to judge. The venues that win the next phase will be the ones where the judging can be trusted as much as the trading, because in the end a prediction market has exactly one job: the trade you win should be the payout you get.
Frequently Asked Questions
How do prediction markets like Polymarket decide who won a bet?
Polymarket settles through UMA’s optimistic oracle: a whitelisted proposer posts an outcome with a bond, anyone can dispute it within a two-hour window by matching the bond, and a disputed market goes to a vote of UMA token holders whose ballots are weighted by staked tokens. Winning shares then redeem for $1 and losing shares for nothing.
Why did Polymarket resolve the Strategy Bitcoin market to No when the sale happened?
The market’s designated source was Strategy’s SEC filings, and the filing that confirmed the 32 Bitcoin sale was published on June 1, one day after the contract expired. Voters and the platform treated the outcome as unconfirmed at expiry, so the market paid No, and two traders sued, arguing that a public-confirmation requirement was applied after the event had already occurred.
What is the difference between how Polymarket and Kalshi resolve markets?
Polymarket uses a decentralized, on-chain oracle where outside token holders can propose and vote on outcomes, while Kalshi determines results in-house against a named data source, with an internal Outcome Review Committee whose decisions are final. One design risks conflicted voters; the other concentrates rule-writing, settlement, and payout inside a single company.
Can AI resolve prediction markets instead of humans?
Partly. UMA’s Optimistic Truth Bot proposes outcomes with roughly 78% accuracy across thousands of markets, reaching over 99% on simple sports and price questions but far lower on open-ended or mention-counting markets. Research suggests AI can auto-settle the clean, high-confidence majority of questions and escalate the ambiguous tail to human review, rather than replacing human judgment entirely.
Which regulator oversees prediction markets in the United States?
Event contracts are the domain of the Commodity Futures Trading Commission, which regulates designated contract markets like Kalshi and Polymarket’s US venue. The Securities and Exchange Commission’s reach here is limited to the tokens that sit alongside the machinery, such as UMA and OLAS, not the event contracts themselves.
By Marcus Okafor, senior markets writer at HOGE Wire, covering the plumbing of on-chain finance.