Interactive paper companion
Two ways accuracy can mislead
Small, manipulable versions of two experiments on prediction-market forecasting. No model or market data is loaded here; the examples isolate the mechanism.
Experiment 01
Delete the quiet moments
A persistence forecast says the next price will equal the last one. On an ordinary tick grid, it is correct whenever the price stays flat. Delete every repeated price and persistence becomes wrong at every remaining step by construction.
Keep every tick and persistence receives credit for correctly predicting quiet steps.
After deduplication, every remaining transition satisfies pₜ ≠ pₜ₋₁. A no-change forecast therefore has exactly 0% direction accuracy.
On a matched population, the same preprocessing inflated persistence RMSE by 27.4%. This toy sequence demonstrates the mechanism, not that empirical estimate.
Experiment 02
Hold accuracy constant
Two forecasts can have the same RMSE and disagree on whether they describe a possible market. Switch between them, then edit the probabilities yourself.
This forecast preserves total probability mass. Its accuracy is not the whole story, but it is internally possible.
Mutually exclusive temperature bins form a distribution. The experiment scores the difference between predicted and realized total mass.
As the points threshold rises, P(total > x) cannot increase. Every upward step is a no-arbitrage violation visible before the outcome resolves.
The common lesson
Evaluation is part of the model.
A benchmark decides what counts as skill. A metric decides what kinds of mistakes are visible. In both papers, changing the evaluation changed the scientific conclusion without changing the forecast.