Interactive paper companion

Two ways accuracy can mislead

Small, manipulable versions of two experiments on prediction-market forecasting. No model or market data is loaded here; the examples isolate the mechanism.

Experiment 01

Delete the quiet moments

A persistence forecast says the next price will equal the last one. On an ordinary tick grid, it is correct whenever the price stays flat. Delete every repeated price and persistence becomes wrong at every remaining step by construction.

Choose a price sequence
Observations16
Flat steps—
Persistence direction accuracy—
Persistence RMSE—

Keep every tick and persistence receives credit for correctly predicting quiet steps.

What is mathematical

After deduplication, every remaining transition satisfies pₜ ≠ pₜ₋₁. A no-change forecast therefore has exactly 0% direction accuracy.

What the paper measured

On a matched population, the same preprocessing inflated persistence RMSE by 27.4%. This toy sequence demonstrates the mechanism, not that empirical estimate.

Experiment 02

Hold accuracy constant

Two forecasts can have the same RMSE and disagree on whether they describe a possible market. Switch between them, then edit the probabilities yourself.

Choose a ladder geometry
Choose a forecast preset
Accuracy 0.020 RMSE
Coherence error 0.000 mass

This forecast preserves total probability mass. Its accuracy is not the whole story, but it is internally possible.

Weather partitions

Mutually exclusive temperature bins form a distribution. The experiment scores the difference between predicted and realized total mass.

Sports CDFs

As the points threshold rises, P(total > x) cannot increase. Every upward step is a no-arbitrage violation visible before the outcome resolves.

The common lesson

Evaluation is part of the model.

A benchmark decides what counts as skill. A metric decides what kinds of mistakes are visible. In both papers, changing the evaluation changed the scientific conclusion without changing the forecast.