h hoge.gg
Subscribe
BTC$67,432.18+2.34%ETH$3,521.44+1.08%SOL$178.62-0.62%BNB$612.30+0.41%XRP$0.6234-0.18%ADA$0.4521+3.12%DOGE$0.1623+1.86%AVAX$38.71-1.24%LINK$17.84+0.92%HOGE$0.00004120+4.21%
BTC$67,432.18+2.34%ETH$3,521.44+1.08%SOL$178.62-0.62%BNB$612.30+0.41%XRP$0.6234-0.18%ADA$0.4521+3.12%DOGE$0.1623+1.86%AVAX$38.71-1.24%LINK$17.84+0.92%HOGE$0.00004120+4.21%
● AI x Crypto

The Forecast Layer: AI Matches the Market, Loses the Bet

Prediction markets became the place AI agents trade, grade each other, and feed signals to other bots. A new benchmark shows the models now match the market on accuracy, yet still cannot beat it at be

For two years the story about prediction markets and artificial intelligence had a single shape: the bots arrive, the bots take the order book, and human traders get priced out. By the autumn of 2026 that part is settled. Autonomous agents now account for more than 30% of the wallets active on Polymarket, and on a typical day 14 of the 20 largest wallets are machines, according to data cited by CoinDesk. The newer and stranger development is what the same markets have become around that trading: the place where AI forecasters are graded, the live probability feed that other software reads to make decisions, and increasingly a system that AI helps run and settle.

Call it the forecast layer. A prediction-market price between zero and one dollar has always been sold as the best available estimate of whether something will happen, from a Federal Reserve cut to a hurricane landfall. That was a claim about crowds. In 2026 the crowd is mostly code, and the question that used to be rhetorical is now an empirical one: can a large language model forecast the future as well as the market it would bet into? A benchmark out of the University of Chicago spent the year answering it, and the result is the most interesting finding the sector produced. The models match the market on accuracy. They still cannot beat it at the only game that pays out.

This piece is about that forecast layer: who builds it, who grades it, whether its numbers can be trusted, and what happens when the forecaster, the trader, the operator, and the judge are all the same kind of system. A note on jurisdiction up front, because it matters later: in the United States these event contracts are the Commodity Futures Trading Commission’s turf, not the SEC’s. The SEC enters only where the tokens and the ordinary securities sit, and, as we will see, the forecast layer is starting to drag those two worlds together.

Two machines that both claim to see the future

It helps to separate two kinds of AI that often get lumped together. The first is the trading agent: software with its own wallet and a model for a brain that reads a market, forms a view, and places bets around the clock. The second is the forecasting model: an LLM asked, cold, to assign a probability to a future event. Both produce a number that is supposed to describe the odds. They are not the same thing, and 2026 made prediction markets the arena where both are measured.

The distinction carries most of the weight in this story. A trading agent is judged on profit and loss. A forecasting model is judged on calibration, whether its stated 70% chances really happen about 70% of the time. You can be excellent at one and poor at the other. A forecaster can be beautifully calibrated and still lose money, because the market has already priced the easy calls and the edge lives in the hard, ambiguous, fast-moving tail. That gap, between being right and getting paid, is the thread that runs through everything below.

Keep that split in mind, because the public conversation blurs it constantly. When a headline says an AI beat a prediction market, it almost always means a forecasting model was better calibrated on a test set, not that a bot made money. When a trader complains that bots ruined the edge, they mean the trading agents, not the forecasters. The two groups even want different things from a market: a forecaster wants clean, objective questions it can reason about, while a trader wants messy, mispriced corners where the crowd is lazy. Most of the confusion in this field comes from treating one as a proxy for the other.

How the bots took the order book

The trading side is the mature part of the market, and it is worth restating only briefly. Prediction markets suit autonomous agents better than almost any other venue in crypto: the action space is a bounded yes or no, the reward is objective and arrives on a schedule, the markets run continuously, and the whole thing is reachable by API. David Minarsch, chief executive of Valory, the company behind the Polystrat agent and the Olas network, described his product to CoinDesk as “an autonomous AI agent that trades on Polymarket 24/7 on behalf of its human user.” Polystrat logged more than 4,200 trades in a single month, with one position returning as much as 376%.

Minarsch is also careful about what the machines cannot do. Simply prompting an off-the-shelf model with a market, he told the same outlet, “usually results in outcomes no better than a coin-flip.” The edge comes from the plumbing around the model, and from picking the right fights: “the long tail of prediction markets is very interesting for AI agents.” Agents hold their funds and sign their trades through the same smart-account machinery that powers the rest of on-chain automation, a setup we walk through in our guide to how smart accounts work in 2026. The headline split is blunt: roughly 37% of agents run a positive book, against something like 7% to 13% of humans.

Prophet Arena: grading the models against the market

If trading agents are judged by their bankroll, forecasting models needed a scoreboard of their own. That is what Prophet Arena provides. Built by a team at the University of Chicago and published under the title LLM-as-a-Prophet, it continuously pulls live questions from regulated prediction markets, mostly Kalshi, asks frontier models to assign probabilities before the events resolve, and then scores them once reality settles the bet. Using live markets is a clever trick: the questions concern genuinely unknown futures, so a model cannot have memorized the answer during training, and the market’s own price serves as the human baseline to beat.

The results are precise and, read carefully, surprising. On calibration the models win. The paper reports that all the tested LLMs showed better calibration than the market baseline, with the strongest posting a calibration error at or below 0.05 against the market’s 0.069. On raw forecasting loss the two are a wash, with model Brier scores clustering around 0.18 to 0.22 versus the market’s 0.187. And yet, on the metric that mirrors actual betting, the models lose. In the authors’ words, “even GPT-5R, the top-ranked model, fails to reach break-even,” with most models returning less than the amount staked. Good enough to forecast, not good enough to win.

BenchmarkWhat it testsGround truthHeadline result
Prophet Arena (LLM-as-a-Prophet, University of Chicago)Forecasting accuracy, the quality of the probabilityLive Kalshi and Polymarket outcomesTop models better calibrated than the market, Brier roughly equal, but every model’s betting return below break-even
Prediction Arena (Arcada Labs)Trading profitability, profit and loss6 frontier models, 10,000 dollars each, 57 daysEvery Kalshi bot lost 16% to 31%, Polymarket roughly breakeven, one model hit a 71% settlement win rate

Calibrated but broke: why accuracy is not alpha

Why would a model that estimates odds as well as the market still bleed money betting into it? Prophet Arena points at one culprit: the models are systematically timid. Across most events, the paper found, the LLMs “consistently output more conservative probabilities” than the market, and even when the crowd is near certain, the models “remain hesitant, rarely producing equally extreme predictions.” A forecaster that refuses to say 97% when 97% is right leaves the most profitable, highest-conviction bets on the table. Calibration keeps you honest; conviction is what pays.

A second benchmark hits the same wall from the trading direction. Prediction Arena, from Arcada Labs, handed six frontier models 10,000 dollars each and let them trade for 57 days early in 2026. On Kalshi every model finished in the red, with losses between 16% and roughly 31%; on Polymarket the field landed near breakeven. The authors’ conclusion, that platform design mattered more than raw model capability, lines up with microstructure work by Philipp Dubach, whose study of Polymarket’s order book found that trade direction inferred from the public feed matches the on-chain truth only around 59% of the time, against roughly 80% for comparable methods on Nasdaq. A bot building signals off that feed can read order flow backwards. The lesson across both papers is the one Minarsch already stated: the model is the easy part.

The market as a signal: wiring odds into your portfolio

If a prediction market is a live probability feed, the obvious next move is to plug it into decisions that have nothing to do with the bet itself. That is exactly what the brokerage Public shipped on 24 September 2026. As Fortune reported, Public launched AI agents that trade event contracts on Kalshi and, more importantly, let a user turn a shifting probability into a trigger for the rest of a portfolio. Tell an agent to buy a stock if the odds of a regulatory approval cross a threshold, to alert you when the market-implied chance of a Fed cut moves, or to adjust an options position as an earnings probability drifts, and the agent watches the feed and acts.

“AI agents can do work for you, and in investing, that means they can monitor markets when you’re not looking at a screen,” said Leif Abraham, Public’s co-founder and co-chief executive. The framing is the tell. The prediction market stops being a casino you visit and becomes an oracle that other machines consume, the way a price feed or a weather API is consumed. It also quietly crosses a regulatory line, because the thing the agent buys off a Kalshi signal might be an ordinary share, which lands back in the SEC’s world. More on that tension below.

When the exchange runs on agents too

The supply side is automating as well. Kalshi has built an AI agent to help run the exchange itself, according to PYMNTS. The tool aggregates relevant news, analyzes what rivals are listing, recommends which contracts the exchange should offer next, and, most consequentially, stress-tests the wording of contracts before they go live. Contract language is where prediction markets quietly break: a bet can hinge on a single ambiguous phrase, and the exchange has already been burned by cases where the words did not match the event, including a Netflix earnings market that turned on a question of pronunciation.

“We actually have an AI engineer in the markets team, where the AI is battle-testing the entire certification, finding out if you go in this direction, maybe there’s a hole here, and all of that,” Kalshi co-founder Luana Lopes Lara said in a Bloomberg interview relayed by PYMNTS. Put the pieces together and a pattern appears: the agent that writes the market, the agents that trade it, and the agent that grades the traders are all models. The only human left in the core loop is the one who decides what to automate next.

That contract-wording job is less glamorous than trading, but it may matter more for the forecast layer’s credibility. A probability is only meaningful if everyone agrees what would make it true, and the most expensive disputes in this market have come from sloppy language rather than bad luck. Handing the first draft of that language to a model that can enumerate edge cases faster than a human lawyer is a real improvement, as long as a human still signs off; it is also one more place where the machine shapes the market before a single bet is placed.

The judge is a model now

Resolution, the step that decides who was right, is the last piece to go autonomous, and it is the most dangerous one to hand to a machine, because a wrong settlement is not a bad trade, it is a broken promise. Polymarket settles through UMA’s optimistic oracle, and UMA has been testing an AI proposer it calls the Optimistic Truth Bot. By the project’s own account, the bot resolves roughly 78% of markets correctly across more than 3,000 tests, climbing above 99% on simple sports and asset-price questions, but it still posts its recommendations to social media rather than on-chain, a tacit admission that it is an assistant, not yet an arbiter.

A cleaner design comes from Andrew Hall, the Davies Family Professor of Political Economy at Stanford and a Hoover senior fellow, writing for a16z crypto. His proposal, laid out in an essay on AI judges, is to fix the resolver at the moment the market is created: the maker specifies “not just the resolution criteria in natural language, but the exact LLM (identified by a timestamped model version) and the exact prompt that will be used to determine the outcome.” Then, as Hall puts it, “the entire resolution mechanism is visible and auditable before anyone places a bet. No rule changes mid-flight, no discretionary judgment calls.” The forecast layer, in other words, is being staffed by machines end to end: listing, trading, grading, and settling.

LayerFunctionThe AI system in 2026Maturity
Listing and operationsChoose and word the contractsKalshi’s markets-team AI (certification stress-testing, listing ideas)In production, supervised
Trading and pricingSet the price by bettingAutonomous agents (Valory’s Polystrat and peers)Dominant, 30%+ of Polymarket wallets
Distribution and signalsTurn odds into actions elsewhereAgentic brokerages (Public’s event-driven agents on Kalshi)Launched September 2026
ResolutionDecide the outcomeUMA’s Optimistic Truth Bot, a16z’s fixed-prompt AI judgesExperimental, advisory
GradingScore the forecastersProphet Arena, Prediction ArenaLive research benchmarks

Markets versus models: which is the better forecaster?

So when a market price and a model’s probability disagree, who should you believe? The honest answer is that they fail in different ways, and the disagreements are becoming routine. Through 2026 the venues and the models openly diverged on the big macro questions, from where Bitcoin finishes the year to whether the Fed cuts or holds, the same event markets that drove so much crypto trading around the election and rate-decision calendar. A market aggregates dispersed private information and the money of people who think they know something; a standalone model aggregates public information and reasoning. The market reacts faster to breaking news, as Prophet Arena found when it watched models lag the price as resolution neared. The model, free of the crowd’s herd behavior, sometimes spots a mispriced long shot the book has ignored.

The synthesis most builders are reaching is that the two are complements, not rivals. A model is a cheap, fast, always-on first opinion; the market is the expensive, incentive-backed second opinion that is dangerous to ignore. The trouble starts when the second opinion is itself mostly models, because then the independence that made the crowd wise begins to erode.

In practice the better desks already run both. A model generates a fast, cheap probability on every open market overnight; the next morning a human or a higher-level agent compares that number to the live price and only acts where the gap is wide and explicable. The model is a screen, the market is the confirmation, and the disagreements are the watchlist. It is a sensible division of labor, and it quietly assumes the market stays the more trustworthy of the two. That assumption is exactly what the monoculture risk puts in question.

The monoculture problem: when the crowd is one model

The wisdom of crowds rests on an assumption that often goes unstated: that the errors of individual forecasters are independent, so they cancel out in the average. A market dominated by autonomous agents strains that assumption, because a large share of those agents are built on a handful of the same frontier models, fed similar data through similar prompts. When one popular model is wrong in a particular way, many agents are wrong together, and the price moves as a herd rather than a crowd. Prophet Arena’s finding that the models cluster toward the same conservative probabilities is a preview of exactly this failure mode at scale.

Correlated agents do more than dull the signal; they invite reflexivity and the order-flow games that plague other automated venues. It is a familiar shape in crypto: a resource gets so abundant and so uniform that its marginal value collapses. The same risk hangs over a forecast layer whose contributors are drawn from one gene pool. Diversity of models, data sources, and strategies is not a nicety here; it is the thing that keeps the forecast worth reading.

A 60 billion dollar September, and the volume you cannot fully trust

A forecast layer is only as credible as the liquidity behind its prices, and here September 2026 sent a mixed signal. Kalshi posted roughly 60 billion dollars in notional volume for the month, a record, even as rivals including Polymarket’s US venue, DKeX, and Novig trimmed its share of American prediction-market volume to about 75%, according to Prediction News. The firm logged its first three-billion-dollar day on 20 September and its biggest week of the year over 14 to 20 September.

The catch is in the word notional. In that record week, DeFiRate counted 15.27 billion dollars of contract volume but only 3.53 billion dollars of actual money changing hands across 85.26 million trades, a reminder that a one-dollar binary contract can be flipped many times for a fraction of a dollar at risk. It gets worse under a microscope. A CNBC analysis found that on 20 September nearly half the dollar volume in Kalshi’s Ether perpetual contracts came from trades sized between 5,495 and 5,505 dollars, a clustering so tight it looks engineered rather than organic. Polymarket, for its part, saw around 2.3 billion dollars actually paid over the month. When the headline number and the at-risk number diverge this far, every downstream consumer of the feed, human or machine, should discount accordingly.

September 2026 metricFigure
Kalshi, notional volume (monthly record)Roughly 60 billion dollars
Kalshi, share of US prediction volumeAbout 75%, eroded by Polymarket US, DKeX, Novig
Kalshi, biggest week (Sep 14 to 20), contract volume15.27 billion dollars
Same week, actual dollars exchanged3.53 billion dollars across 85.26 million trades
Kalshi, first three-billion-dollar day (Sep 20)3.17 billion dollars
Polymarket, dollars actually paid over the monthAround 2.3 billion dollars

Who regulates a forecast?

The legal picture is unsettled in a way that the forecast-layer framing actually sharpens. Event contracts themselves sit with the Commodity Futures Trading Commission, which proposed a rewrite of its Rule 40.11 on event contracts in June 2026 and is still working through the comments. The courts, meanwhile, are split: as a Congressional Research Service brief lays out, one federal appeals circuit has treated Kalshi’s sports contracts as swaps under federal law, while others have sided with states that call them sports bets, and New Jersey has asked the Supreme Court to step in. No final decision had landed by early October.

The forecast layer adds a wrinkle the circuit split does not capture. When an agent reads a Kalshi probability and uses it to buy a stock, as Public’s product allows, the trade that results is in an ordinary security, which is the SEC’s domain, not the CFTC’s. A probability generated under commodities law becomes an input to a securities transaction, with no single regulator overseeing the handoff. The agencies have spent 2026 in an uneasy truce over crypto generally, a detente we traced in our report on SEC enforcement after the crackdown; the moment prediction-market signals start steering regulated portfolios, that truce gets a new and awkward test case.

For anyone building on top of this, the practical takeaway is to treat a prediction-market probability as a regulated input, not a free one. The venue that produced it, the contract’s exact wording, and the jurisdiction of whoever acts on it all travel with the number. A compliance team that has signed off on an agent trading event contracts has not necessarily signed off on that same agent rotating a stock portfolio on the back of them, and the gap between those two approvals is where the next enforcement surprise probably lives.

The blind spots: recall, timing, and the long tail

For all the talk of machine forecasters matching the crowd, the Prophet Arena researchers were blunt about where the models still fail, and the failures are instructive. The models struggle with precise event recall, muddling time-stamped facts in domains like weather and politics even when they have the general picture right. They misjudge sources, sometimes getting less accurate when more information is added because they cannot tell a reliable input from a noisy one. And they aggregate breaking news slowly, falling behind the market exactly when an event is about to resolve and the last-minute information matters most.

Those are forecasting errors. The feed also carries settlement errors, which are worse, because a feed is an oracle and oracles get gamed. The lesson that oracle manipulation went multi-chain this year is the relevant one: if the output of the forecast layer becomes an input to real money elsewhere, the incentive to corrupt that output grows with every dollar that depends on it. A probability is only as trustworthy as the mechanism that will eventually declare it true or false, which is why the resolution question refuses to go away no matter how good the forecasting gets.

What a trustworthy forecast layer would need

If prediction markets are going to serve as infrastructure that other systems depend on, rather than a venue people visit, the bar rises. A few requirements follow directly from the problems above.

  • Provenance. A consumer of a probability should be able to tell whether it was set by deep human conviction, by a herd of correlated agents, or by wash volume dressed up as liquidity.
  • Auditability. Resolution rules, including the exact model and prompt when an AI judge is used, should be fixed and visible before trading opens, as Hall’s design insists.
  • Calibration tracking. Public scoreboards like Prophet Arena should be the norm, so a feed’s historical reliability is a known quantity rather than a marketing claim.
  • Separation of forecasting from trading. A model that grades outcomes should not be run by a party with a position in them, the conflict that has long dogged on-chain resolution.
  • Human escalation. The long tail of ambiguous, novel, and fast-moving events is where machines are weakest and should be routed to people.

None of this is exotic. It is the ordinary discipline of any data product that other people build on, applied to a market that grew up as entertainment and is now being asked to behave like a utility.

The truth machine, mostly machines now

Prediction markets were always pitched as a truth machine, a way to turn scattered belief into a single honest number. In 2026 that machine became, to a startling degree, machines: agents set the prices, a brokerage’s agents read them, an exchange’s agent writes and tests the contracts, and an oracle’s agent proposes the result, while academic benchmarks keep score. The remarkable part is how well it works. Autonomous forecasting has found, in these markets, the one place where its reward function is legible enough to make the whole loop pay.

The sobering part is what the benchmarks actually show. The models match the market on accuracy and still lose money, because calibration is not conviction, and the market’s edge is the diversity and the incentives of the people, and now the machines, inside it. As that diversity narrows toward a few shared models, the edge is exactly what is at risk. The companies building this layer are worth watching less for their token prices, which remain tiny micro-caps where framework quality and market value parted ways long ago, a split we traced in Gensyn’s low-float $AI, than for whether they can keep the forecast layer honest while handing more of it to the machines. On that question the jury, human or otherwise, is still out.

Frequently Asked Questions

Can AI really forecast the future as well as prediction markets?

On accuracy, close to it. The Prophet Arena benchmark found that top models are actually better calibrated than market prices and roughly tied on raw forecasting error, but they are systematically too conservative, so when they bet into the market they still fail to turn a profit. The short version is that AI matches the market as a forecaster and loses to it as a bettor.

What is a prediction-market agent?

It is autonomous software, usually a large language model paired with its own crypto wallet, that reads a market such as Kalshi or Polymarket, forms a probability view, and places and manages bets without a human clicking each time. Agents now make up more than 30% of active Polymarket wallets and most of the largest ones.

Are prediction markets regulated by the SEC?

Not directly. In the United States event contracts fall under the Commodity Futures Trading Commission, which is rewriting its event-contract rules while courts and states fight over sports contracts. The SEC enters only around the tokens and around ordinary securities, for example when an AI agent uses a prediction-market probability as a trigger to buy a stock.

Why do AI trading bots lose money even when they are accurate?

Because being well-calibrated is not the same as having an edge. The market has already priced the easy, obvious outcomes, so profit lives in the hard, ambiguous, fast-moving tail, and that is where models are weakest. Benchmarks found most AI traders finished in the red on Kalshi, and research on Polymarket’s order book showed that signals read off the public feed can get order flow backwards.

What does it mean to call prediction markets a forecast layer?

It means the market is no longer just a place people go to bet; its live probabilities are becoming a data feed that other software reads and acts on, from brokerages that wire the odds into portfolio decisions to benchmarks that use the outcomes to grade AI models. The market turns into shared infrastructure, which raises the bar for how trustworthy its numbers have to be.

By Marcus Okafor, HOGE Wire staff.

Share 𝕏 Post Telegram