Methods & Results in Depth
A figure-led walk through the data, the supervised efficiency baselines, the unsupervised latent-state model selection, the absorption-aware latent-intensity model (MMPP + variational inference), the non-parametric Bayesian latent-types flagship, and the inference validation — with the methods, the numbers, and the honest nulls and self-corrections that the high-level overview compresses away.
Juan Mediavilla · University College London · MSc Machine Learning · Academic Supervisor: Prof. Philip Treleaven · Technical Advisor: TBC
Orientation
The thesis advances by cheap, falsifiable, pre-registered probes, each with named gates and a committed verdict — including the verdicts that came back negative. The arc is a progression of machine-learning methods: supervised baselines rule out the easy story, unsupervised latent-state learning identifies the real structure, a latent-variable generative model is built and inferred by variational methods, then validated to the standard that retracted earlier candidate findings.
Supervised directional classifiers + Brier calibration across 1,348 markets. Five no-edge results redirect the thesis from prediction to learning price formation.
Falsify the variance-regime HMM; learn that persistence lives in a latent repricing-intensity state, not per-bin variance.
A model ladder M0→M1→M2: homogeneous Poisson, then a latent hot/cold MMPP, then absorption-aware modulation by time-to-resolution, fitted by variational inference.
Held-out likelihood, hierarchical Bayesian pooling, a non-parametric Bayesian latent flagship, and a walk-forward causal test that bounds the predictive claim.
Every number on this page is drawn from a committed result; the originating probe is named in light grey beside each figure (e.g. market_census_v0, latent_types_v0) so a reader can trace it. Confidence intervals are market-clustered bootstrap unless stated.
1 · Data
No public dataset of a prediction venue's live order book existed. The thesis builds one: a gap-audited, millisecond-resolution capture of the most active Polymarket markets — full order-book deltas, periodic snapshots, the trade tape and resolutions — reconstructed to a mid-price path by delta replay anchored on snapshots, captured continuously through each market's settlement.
Honest limits carried throughout: a finite collection window, liquidity-based selection of the top markets, and single-venue scope. Sample sizes are stated at every claim — the dynamics probes use 10 liquid markets, the hierarchical backbone 34, the absorption gold cohort 9, the freeze-at-scale cohort 93, and the latent flagship 52 book-covered markets.
2 · Supervised baselines
Before modelling dynamics, the thesis establishes why dynamics is the right object: the price is already informative, so the productive question is how it forms, not how to beat it. A battery of pre-registered, fee-aware tests finds no out-of-sample directional edge — a load-bearing negative result.
On the binary markets observed through resolution, the late market price scores a Brier of 0.0585 against the realised 0/1 outcome — far below the 0.166 a base-rate prior achieves on the same set (71 binary markets, structure_report_v0). The market beats every coarse prior tested, in every category.
Four directional strategies (mean-reversion, longshot-fade, depth-imbalance, regime-conditioned) and a cross-market lead-lag test all fail out of sample under multiple-testing correction. The lead-lag screen is the sharpest: of 192 directional propagation tests across market pairs, 0 survive a Benjamini–Hochberg correction (leadlag_probe_v0) — the out-of-sample incremental R² of the lagged term is at best zero. There is no riskless arbitrage either: across binary complements the combined ask never falls below $1.001, and the gross ~3¢ mispricings that do appear in 1X2 partitions are erased by fees (~2¢ over three legs) and survive for only 36 ms — 0 are tradeable net (arb_scan_v0).
The same conclusion recurs across directional, dependence, and relative-value tests run later in the project: where a gross edge appears, the two-legged cost stack (spread plus fee) erases it. The finding is robust to removing the largest markets and to split-half resampling.
3 · Unsupervised model selection
Two pre-registered probes adjudicate between competing models of the price path. The natural model — a regime-switching model of price variance — is decisively falsified. The persistent structure instead lives in the intensity of repricing: how often the market moves.
Direct tests bear out the reframe: repricing events cluster strongly and are non-Poisson on all ten markets (coefficient of variation 2.2–5.3), the rate autocorrelation persists to ~50 min, and a market's first-half and second-half repricing rates correlate at ρ=0.74 — hot stays hot. A sticky-variance model offered the persistence a fair chance and still could not recover it (only 2/10 markets reached a 5-minute dwell, and both were prior-driven with no likelihood support). The conclusion dictates the model class: the price is sticky with clustered repricing, and the latent state to model is the event intensity.
4 · The latent-variable model
The model is built as a ladder, each rung forced by the evidence of the rung below. A homogeneous Poisson baseline is rejected by the clustering result; a Markov-modulated Poisson process introduces a latent hot/cold intensity state with genuine dwell; and the contribution model makes the intensity a function of time-to-resolution τ — the defining feature of a market that is absorbed at $0 or $1 by a known deadline.
Constant repricing rate. Rejected outright by Chapter 3's clustering: the MMPP beats it on held-out predictive likelihood on every market tested.
A latent hot/cold state with genuine dwell. The cold state is near-silent and the hot state bursty; episodes persist for minutes to hours rather than flickering.
Intensity and state transitions become functions of time-to-resolution. The absorbing variant (M2b) — repricing decays to zero as the market locks onto its answer — is preferred by held-out likelihood over a free time-varying form (M2a).
Per-market fits are pooled into a population model, θm ~ N(μ,Σ), by partial pooling (Metropolis-within-Gibbs). Pooling helps every one of the 34 markets (held-out gain +11,400 nats, CI [3,256, 23,292]); the rate/dwell block converges cleanly (R̂≈1.00) while the trigger block is only weakly identified (R̂ 1.23–1.37) — an early signal of the ceiling the flagship later measures. The population covariance is low-rank: its top three eigen-axes carry 87.6% of the variance (0.63 / 0.15 / 0.09).
5 · The freeze
The absorption hypothesis — does repricing change systematically as a market approaches its deadline? — has a clean, and surprising, answer. Most markets freeze: repricing collapses as the price locks onto its outcome. But the regime is path-dependent, and a self-correction along the way reshaped the claim.
6 · Non-parametric Bayes
The flagship chapter asks whether the consensus-versus-calendar trigger axis becomes a learnable dimension that separates types of market — fit nonparametrically as a Dirichlet-process mixture over per-market parameters by amortized variational inference, on the 52 book-covered markets. The pre-registration committed three outcomes, including the null and an explicit promise not to salvage a weak positive into a false discovery.
The broader latent structure is real but weak and continuum-like. The nonparametric prior, free to collapse to one type, instead settled on roughly four soft, overlapping modes in the rate/intensity block — with a non-monotonic held-out curve (K=2,3 are worse than K=1) that is the signature of weak rather than crisply-separated structure. The strict structure gate is a narrow miss (CI [−0.07, +0.64], includes 0). Crucially, the latent representation still beats observed context — it carries +10.3 nats [5.5, 15.8] over the binary regime label and more over category — so a hierarchical backbone plus a soft low-k mixture remains a defensible descriptive spine, provided the trigger axis is reported as unresolved.
7 · Inference validation
A walk-forward, causal-filter test subjects the model to the standard that retracted earlier candidate findings — and bounds the predictive claim honestly. The latent state is descriptively rich, but it does not forecast.
This is the disciplined boundary of the whole project: the MMPP is a faithful descriptive and detection model of repricing intensity, and the amortized VI is a faithful inference method — but neither is a predictive claim. The honest characterisation, an absorption-aware intensity model whose latent state is real, detectable, and bounded by an identification ceiling, is the thesis's result whichever way the central hypothesis fell.
8 · Methodology
The methodological spine is the part a technical reader should weigh most heavily. Every probe is pre-registered with named gates and a committed verdict; inference is checked against simulation; confidence intervals are market-clustered bootstrap; and findings were retracted when the evidence turned. The project's credibility rests on the negative results as much as the positive ones.
| Probe | Question | Key number | Verdict |
|---|---|---|---|
| Efficiency | Does the price beat a base-rate prior? | Brier 0.0585 vs 0.166 | efficient |
| Directional edge | Any out-of-sample directional edge? | 0 / 192 lead-lag | no edge |
| Variance regime | Do volatility regimes persist? | 0 / 10 markets | falsified |
| Intensity clustering | Is repricing non-Poisson and clustered? | 10 / 10, CV 2.2–5.3 | established |
| MMPP (M1) | Latent hot/cold state with dwell? | beats M0 10/10 | established |
| Hierarchical pool | Does partial pooling help? | +11,400 nats, 34/34 | go |
| Freeze (M2) | Does intensity change as τ→0? | rate-ratio 0.15 / spike 5.77× | two regimes |
| Latent trigger axis | Do trigger types separate? | SNR 0.22, p=0.66 | bounded null |
| Walk-forward | Does the latent state forecast? | ties Hawkes | detect ≠ anticipate |
All figures and numbers on this page are drawn from committed results in the project repository; the originating probe is named beside each. The companion chapter framework and slide deck give the higher-level views.