The forecast does not beat the market. It tells you when to sit out.
The obvious way to trade a temperature market is to fetch a weather model, read tomorrow's high, and buy that bucket. We spent a long time trying to make that work and it does not. A well-run forecast lands on the same answer the market already has, most of the time, and when the two disagree the model is usually the one that is wrong. The useful finding is smaller and stranger: the forecast is close to worthless as a predictor and quite valuable as a veto.
What a temperature market looks like from the inside
A Polymarket temperature event is eleven contracts on one ladder. Ten of them are single whole degrees, the two on the ends are open ranges, and exactly one settles at a dollar. Resolution comes from one airport weather station, rounded to whole degrees Celsius, for the local calendar day. There is no partial credit. Twenty seven point four is twenty seven, and a contract on twenty eight pays nothing.
Because there are eleven rungs and only one winner, the market never gets very confident. Across a validation sample the leading bucket sat around forty three cents a day and a half before close. That is a market saying it thinks it knows, but not very hard.
The naive trade, priced honestly
Start with the simplest rule available: at a fixed hour before close, buy whichever bucket has the highest bid. No model, no view, just deference to the crowd. Across our sample that trade wins about forty four percent of the time at an average price near forty one cents. After the spread you actually pay, it returns somewhere in the high single digits. Positive, but thin enough that a bad month erases it and a wide market eats it alive.
That number is the honest baseline, and it is worth sitting with. The crowd is already good. Anything you add has to beat a market that is close to fairly priced.
Adding a forecast makes it worse, not better
The tempting next step is to trust your own model over the price. We ran that. Take an ECMWF run, correct it for the station, round it, buy that bucket regardless of what the market thinks. The result was not a small edge in the wrong direction, it was closer to nothing at all. On the subset of days where our corrected forecast pointed at a different bucket than the market favorite, returns were statistically indistinguishable from zero.
Read that carefully, because it is the whole article. Those are exactly the days where you have a view. They are the days you would feel clever. And they are worth nothing.
The version that works is subtraction
Flip the logic. Do not use the forecast to pick. Use it to decide whether to participate at all.
The rule becomes: compute your corrected forecast, look at the market favorite, and if they name the same bucket, buy the market favorite. If they disagree, do nothing that day in that city. You never trade against the market. You only trade alongside it, and only when your own model independently arrives at the same place.
About forty three percent of city-days pass that test. On those days the same buy that returned high single digits unfiltered returns roughly two and a half times more, and the statistical confidence roughly doubles. The trade did not get better. The bad half of it got removed.
Why agreement carries information
A weather model and a prediction market are close to independent instruments. The model knows atmospheric physics and nothing about traders. The market knows whatever its participants collectively know, which includes other models, local knowledge, and the fact that someone is watching the observations roll in. When two roughly independent estimates land on the same integer, the odds that the integer is right go up in a way neither one alone justifies.
When they split, you have learned something too, just not something tradeable. You have learned that at least one of your two instruments is wrong by a degree, and you have no reliable way to tell which. Sitting out is not timidity, it is the correct response to a genuine tie.
The station correction is not optional
One detail decides whether any of this functions. A grid-point forecast is not a station reading. The model resolves a square of atmosphere, the market resolves one thermometer at one airport, and the gap between them is stable and often large.
Measured against settlement over a validation window, correcting for that offset moved bucket accuracy from twenty three and a half percent to thirty two percent, and cut mean absolute error from about one point four degrees to about one point one. In the cities with the largest offsets the effect was dramatic. One airport ran more than two degrees below its grid point consistently, and uncorrected the raw forecast named the right bucket less than one time in ten.
That is not a tuning parameter, it is a units problem. Ten of fourteen cities improved, three got slightly worse, and the ones that got worse were the ones whose offset was already near zero. If you skip this step the filter fires on the wrong bucket and you have built an expensive random number generator.
Accuracy is not the same as profit
A result that surprised us: as close approaches, the market gets very accurate and the trade gets worse.
Two hours before close, the leading bucket wins about seventy seven percent of the time. That is an excellent forecast. It also costs seventy three cents, so being right pays twenty seven. Run the numbers and the return is worse than it was a day and a half earlier, when the favorite was wrong more than half the time but cost forty three cents.
Hit rate is a vanity metric in a market that prices. What you are paid for is the distance between what you know and what the price already reflects, and that distance closes long before the uncertainty does.
What this does not tell you
The sample behind these numbers is a few hundred filtered trades across fourteen cities and roughly two months of validation. That is enough to see the direction and not enough to trust any single city's figure. Per-city samples run in the teens and twenties, where a five percentage point swing is noise.
It also assumes you are paying the spread honestly. Backtests on price history quietly assume you transact at the last trade, which is not a thing you can do. Once real bid-ask costs are subtracted, several apparently attractive entry times stop being attractive. And market structure can change. A crowd that gets better makes this filter narrower, not wider.
The transferable idea
Nothing here is really about weather. It is about what to do with a model when the thing you are modelling already has a price on it.
The instinct is to look for disagreement, because disagreement is where the profit story lives. In a market that is roughly efficient, disagreement is mostly where your error lives. Agreement between independent estimates is the rarer and more useful signal, and the discipline it buys you is the discipline to trade less.
Frequently asked
Does this work on any weather market? The mechanism should generalize wherever settlement is a single objective reading and the market is liquid enough to have a real favorite. It will not survive markets where the ladder is thin or one rung dominates.
Which model? We used a single deterministic run rather than an ensemble, mostly for reproducibility. An ensemble would likely sharpen the filter, since agreement across members is itself information.
Is the filter just picking easy days? Partly, and that is fine. The filtered days do have a slightly higher-priced favorite, meaning the market was already more confident. The return improves anyway, which means the filter is finding more than just confidence.
Where to read these markets
Temperature ladders update continuously and the interesting movement happens when a new model run publishes, not on a schedule that suits a human. If you want to watch them live, SmartX carries the same Polymarket books with the ladder laid out rung by rung.