Work in progress

This page is actively being worked on. Apologies for the figures — some were poorly generated and have overlapping text, along with a few other small errors still to be cleaned up. The data, methods and results behind them are real; the presentation is catching up. I update the site as I hit big milestones.

Methods & Results in Depth

The machine learning behind the thesis

A figure-led walk through the data, the supervised efficiency baselines, the unsupervised latent-state model selection, the absorption-aware latent-intensity model (MMPP + variational inference), the non-parametric Bayesian latent-types flagship, and the inference validation — with the methods, the numbers, and the honest nulls and self-corrections that the high-level overview compresses away.

Juan Mediavilla · University College London · MSc Machine Learning · Academic Supervisor: Prof. Philip Treleaven · Technical Advisor: TBC

Orientation

A probe-and-gate arc

The thesis advances by cheap, falsifiable, pre-registered probes, each with named gates and a committed verdict — including the verdicts that came back negative. The arc is a progression of machine-learning methods: supervised baselines rule out the easy story, unsupervised latent-state learning identifies the real structure, a latent-variable generative model is built and inferred by variational methods, then validated to the standard that retracted earlier candidate findings.

Probe I

Supervised baselines

Supervised directional classifiers + Brier calibration across 1,348 markets. Five no-edge results redirect the thesis from prediction to learning price formation.

Probe II

Unsupervised model selection

Falsify the variance-regime HMM; learn that persistence lives in a latent repricing-intensity state, not per-bin variance.

Model

Latent-variable model + VI

A model ladder M0→M1→M2: homogeneous Poisson, then a latent hot/cold MMPP, then absorption-aware modulation by time-to-resolution, fitted by variational inference.

Validate

Inference validity

Held-out likelihood, hierarchical Bayesian pooling, a non-parametric Bayesian latent flagship, and a walk-forward causal test that bounds the predictive claim.

Every number on this page is drawn from a committed result; the originating probe is named in light grey beside each figure (e.g. market_census_v0, latent_types_v0) so a reader can trace it. Confidence intervals are market-clustered bootstrap unless stated.

1 · Data

A high-frequency prediction-market corpus

No public dataset of a prediction venue's live order book existed. The thesis builds one: a gap-audited, millisecond-resolution capture of the most active Polymarket markets — full order-book deltas, periodic snapshots, the trade tape and resolutions — reconstructed to a mid-price path by delta replay anchored on snapshots, captured continuously through each market's settlement.

250
most-active markets streamed at millisecond order-book resolution
1,348
resolved markets, ~$445M notional, in the efficiency corpus
~23 min
total audited feed-gap over the clean four-day server window
masked
gaps are recorded and never interpolated — loss is known, not silent
Order-book days collected, and the markets each finding rests on
Fig 1The data foundation. Top: the order-book days collected in 2026 across two windows — an early laptop corpus and four clean server days — with masked feed gaps marked rather than hidden. Bottom: how many markets each downstream claim rests on, so the small-n findings (the absorption gold cohort, the latent flagship) are visible as such. — dynamics_probe_v0, absorption_v1, market_census_v0

Honest limits carried throughout: a finite collection window, liquidity-based selection of the top markets, and single-venue scope. Sample sizes are stated at every claim — the dynamics probes use 10 liquid markets, the hierarchical backbone 34, the absorption gold cohort 9, the freeze-at-scale cohort 93, and the latent flagship 52 book-covered markets.

2 · Supervised baselines

Supervised baselines & calibration: the venue is near-efficient

Before modelling dynamics, the thesis establishes why dynamics is the right object: the price is already informative, so the productive question is how it forms, not how to beat it. A battery of pre-registered, fee-aware tests finds no out-of-sample directional edge — a load-bearing negative result.

On the binary markets observed through resolution, the late market price scores a Brier of 0.0585 against the realised 0/1 outcome — far below the 0.166 a base-rate prior achieves on the same set (71 binary markets, structure_report_v0). The market beats every coarse prior tested, in every category.

Late market-price Brier far below the base-rate prior in every category
Fig 2The venue is efficient. Base-rate Brier (grey) versus late-price Brier (teal) by category, binary markets in the pre-fee historical corpus. The late price scores far below the uninformed base rate everywhere — the market is well-calibrated against realised outcomes. — market_census_v0

Five independent no-edge results

Four directional strategies (mean-reversion, longshot-fade, depth-imbalance, regime-conditioned) and a cross-market lead-lag test all fail out of sample under multiple-testing correction. The lead-lag screen is the sharpest: of 192 directional propagation tests across market pairs, 0 survive a Benjamini–Hochberg correction (leadlag_probe_v0) — the out-of-sample incremental of the lagged term is at best zero. There is no riskless arbitrage either: across binary complements the combined ask never falls below $1.001, and the gross ~3¢ mispricings that do appear in 1X2 partitions are erased by fees (~2¢ over three legs) and survive for only 36 ms0 are tradeable net (arb_scan_v0).

Two profit probes, both erased by the cost stack
Fig 3Two ways the cost stack wins. Left — a longshot decay-harvest earns a gross premium of +0.49¢/contract but a conservative 1¢ half-spread turns it to −0.53¢ net; the spread is the killer. Right — fading cold-regime jumps reverts a real but tiny amount, and the fee alone flips +1.6¢ gross to −0.3¢ before spread is applied; the fee is the killer. Inefficiencies exist but are smaller than the cost to capture them — itself evidence of efficiency. — longshot_decay_v0, jump_reversion_v0

The same conclusion recurs across directional, dependence, and relative-value tests run later in the project: where a gross edge appears, the two-legged cost stack (spread plus fee) erases it. The finding is robust to removing the largest markets and to split-half resampling.

3 · Unsupervised model selection

Latent-state learning: persistence lives in repricing intensity, not variance

Two pre-registered probes adjudicate between competing models of the price path. The natural model — a regime-switching model of price variance — is decisively falsified. The persistent structure instead lives in the intensity of repricing: how often the market moves.

Reconstructed YES-price paths: sticky-flat punctuated by clustered bursts
Fig 4What a reconstructed path looks like. Three markets, YES-price on the 10-second grid: long sticky-flat stretches punctuated by clustered bursts (shaded where a 30-min window holds ≥3 repricings). Resolving markets go silent — the price freezes onto its answer — before settlement. — absorption_v1
Variance-regime HMM flickers; dwell is one bin at every grid
Fig 5Why the variance model failed. The "active" variance-regime bins never form sustained phases — they flicker one bin at a time, and the median dwell is ≈1 bin at 1 s, 10 s and 60 s alike. Coarsening the grid does not manufacture persistence (0/10 markets). — dynamics_probe_v0/v1
Repricing-count autocorrelation positive to ~50 min; gaps over-dispersed
Fig 6The clustering evidence. Left: autocorrelation of the per-minute repricing count stays positive out to ~50 min — busy minutes follow busy minutes. Right: inter-event gaps sit well above the exponential line (CV ≈4), long quiet stretches punctuated by tight bursts. Non-Poisson on 10/10 markets. — dynamics_probe_v1

Direct tests bear out the reframe: repricing events cluster strongly and are non-Poisson on all ten markets (coefficient of variation 2.2–5.3), the rate autocorrelation persists to ~50 min, and a market's first-half and second-half repricing rates correlate at ρ=0.74 — hot stays hot. A sticky-variance model offered the persistence a fair chance and still could not recover it (only 2/10 markets reached a 5-minute dwell, and both were prior-driven with no likelihood support). The conclusion dictates the model class: the price is sticky with clustered repricing, and the latent state to model is the event intensity.

4 · The latent-variable model

An absorption-aware latent-intensity model, fitted by variational inference

The model is built as a ladder, each rung forced by the evidence of the rung below. A homogeneous Poisson baseline is rejected by the clustering result; a Markov-modulated Poisson process introduces a latent hot/cold intensity state with genuine dwell; and the contribution model makes the intensity a function of time-to-resolution τ — the defining feature of a market that is absorbed at $0 or $1 by a known deadline.

M0

Homogeneous Poisson — rejected

Constant repricing rate. Rejected outright by Chapter 3's clustering: the MMPP beats it on held-out predictive likelihood on every market tested.

verdict rejected 10/10 · CV 2.2–5.3 > 1 · dynamics_probe_v2
↓  add a latent state
M1

Markov-modulated Poisson — established

A latent hot/cold state with genuine dwell. The cold state is near-silent and the hot state bursty; episodes persist for minutes to hours rather than flickering.

λcold 0.041/min · λhot 1.141/min (population means) · hot dwell 2–139 min · dynamics_probe_v2, hier_mmpp_v0
↓  modulate by time-to-resolution τ
M2

Absorption-aware — the contribution

Intensity and state transitions become functions of time-to-resolution. The absorbing variant (M2b) — repricing decays to zero as the market locks onto its answer — is preferred by held-out likelihood over a free time-varying form (M2a).

M2b preferred · AIC 121,642 vs 123,344 · trigger is price/consensus-driven · m2_session2/3
MMPP fit: event raster, P(hot) heat strip, fitted rate stepping cold to hot
Fig 7The established model, in one picture. Three aligned views of a single MMPP fit to one active market over its densest 16-hour window: the event raster, the inferred P(hot) heat strip, and the fitted rate λ(t) stepping between a near-silent cold state (λ≈0.03/min) and bursty hot states (λ up to ≈2.9/min). The model tracks the market's moods. — dynamics_probe_v2

A hierarchical backbone and variational inference

Per-market fits are pooled into a population model, θm ~ N(μ,Σ), by partial pooling (Metropolis-within-Gibbs). Pooling helps every one of the 34 markets (held-out gain +11,400 nats, CI [3,256, 23,292]); the rate/dwell block converges cleanly (R̂≈1.00) while the trigger block is only weakly identified (R̂ 1.23–1.37) — an early signal of the ceiling the flagship later measures. The population covariance is low-rank: its top three eigen-axes carry 87.6% of the variance (0.63 / 0.15 / 0.09).

Population parameter correlation matrix; low-rank structure
Fig 8Population structure. The cross-parameter correlation matrix Σ; its leading three eigen-axes account for 87.6% of between-market variation — markets vary along a few shared directions, not independently. — hier_mmpp_v0
Amortized VI vs EM/MCMC sanity: ELBO matches, near-zero amortization gap
Fig 9Inference sanity. The flagship is fit by amortized variational inference: the ELBO (−479.2) matches the EM data-loglik (−478.8), the amortization gap is 2×10−4, and VI and EM recover the same assignments (ARI 0.85). A faithful inference network, not a lossy shortcut. — latent_types_v0

5 · The freeze

Two regimes at the boundary, and what triggers them

The absorption hypothesis — does repricing change systematically as a market approaches its deadline? — has a clean, and surprising, answer. Most markets freeze: repricing collapses as the price locks onto its outcome. But the regime is path-dependent, and a self-correction along the way reshaped the claim.

Repricing rate collapses as time-to-resolution approaches zero
Fig 10The freeze. Left: repricing rate versus time-to-resolution (log axis, →0 = resolution) for the 9 gold markets and their pooled mean — intensity holds near baseline, then collapses about 3.2 h out (onset IQR 127–429 min). The pooled final-hour/baseline rate-ratio is 0.15, 95% CI [0.02, 0.59]. Right: the fitted MMPP hot-state occupancy falls to ≈0 in the final hour on every market tested. — absorption_v1
Self-correction — "venue-wide freeze" was wrong. The early reading that the freeze is universal did not survive scale. At n=93 the cohort splits into two regimes: markets whose answer is revealed before the wire freeze (pooled rate-ratio 0.39 [0.26, 0.59]), while markets resolved by an event at the wire — a match ending, a number printing — spike. The at-wire spike class repriced at 5.77× baseline [4.03, 7.98] on 148 markets and replicated out of sample (test-half 7.56, design-half 4.67). The freeze is not a calendar effect; it is driven by when consensus is reached.

Per-market final-hour/baseline rate-ratio with bootstrap CIs
Fig 11Freeze heterogeneity. Per-market final-hour/baseline rate-ratio (log axis; <1 = froze) with clustered-bootstrap CIs, by category. Scheduled/threshold markets freeze hard; the two geopolitics markets split — one freezes, one reprices to the wire — flagged as open at n=2. — absorption_v1
Trigger coefficients: price/consensus-driven, calendar component in book
Fig 12What triggers the regime. The fitted trigger is consensus/price-driven (βprice +0.82 [0.76, 0.88]); a real calendar component (βτ) appears only when the order book is used as the instrument (+0.66), correcting an earlier "calendar ≈ 0" reading. — m2_session2/3

6 · Non-parametric Bayes

Dirichlet-process latent types (amortized VI) — a pre-registered bounded null

The flagship chapter asks whether the consensus-versus-calendar trigger axis becomes a learnable dimension that separates types of market — fit nonparametrically as a Dirichlet-process mixture over per-market parameters by amortized variational inference, on the 52 book-covered markets. The pre-registration committed three outcomes, including the null and an explicit promise not to salvage a weak positive into a false discovery.

The headline is the null. The trigger axis does not become identifiable at n=52. Between-type signal-to-noise on the trigger slopes is 0.22 (below 1), a permutation test cannot distinguish trigger-type separation from noise (p=0.66), and the one positive signal — a held-out lift of +0.62 — is a marginal-density artifact, not a relabelling. This is the fourth independent corroboration of an n≈50 ceiling seen earlier in the hierarchical backbone and the m2 sessions. The result is a measurement fact about 52 markets, not a modelling failure — and it is reported as a first-class finding.

Per-coordinate signal-to-noise: rate coords > 1, trigger coords < 1
Fig 13The ceiling, measured. Per-coordinate between-market signal divided by measurement noise. The rate/dwell coordinates exceed 1 (they can be estimated per market, so structure can live there); the trigger slopes βτ and βprice fall below 1 — noise exceeds signal, so a latent layer has nothing crisp to cluster on. — latent_types_v0
Trigger-type separation indistinguishable from a permutation null
Fig 14Trigger identification fails. The load-bearing gate: trigger-type separation (SNR 0.22) sits inside the permutation null (p=0.66). Two of the three pre-registered conditions fail decisively; the verdict is not identified. — latent_types_v0

What the data does support: a soft continuum

The broader latent structure is real but weak and continuum-like. The nonparametric prior, free to collapse to one type, instead settled on roughly four soft, overlapping modes in the rate/intensity block — with a non-monotonic held-out curve (K=2,3 are worse than K=1) that is the signature of weak rather than crisply-separated structure. The strict structure gate is a narrow miss (CI [−0.07, +0.64], includes 0). Crucially, the latent representation still beats observed context — it carries +10.3 nats [5.5, 15.8] over the binary regime label and more over category — so a hierarchical backbone plus a soft low-k mixture remains a defensible descriptive spine, provided the trigger axis is reported as unresolved.

Posterior over number of types: ~4 soft modes, decaying weights
Fig 15How many types? The held-out-optimal count is k≈4 (bootstrap mean 4.09), but the in-sample weight spectrum decays smoothly (0.38 / 0.19 / 0.11 / 0.10 / …) — one dominant mode and a tail, a continuum rather than a clean k=3. — latent_types_v0
Soft market types in theta-space: weak, overlapping modes
Fig 16Soft types in θ-space. Markets coloured by their dominant soft type. The modes overlap heavily and do not reduce to category or regime — a continuum of liquid-active markets, not crisp clusters. Type assignments replicate at split-half ARI 0.70. — latent_types_v0

7 · Inference validation

Walk-forward validation: the model detects; it does not anticipate

A walk-forward, causal-filter test subjects the model to the standard that retracted earlier candidate findings — and bounds the predictive claim honestly. The latent state is descriptively rich, but it does not forecast.

Causal filter detects hot states quickly but does not lead the data
Fig 17Detection, not anticipation. The causal filter reliably and quickly identifies hot states (detection score 0.68 per market → 1.00 pooled), but it does not precede the data: apparent "leads" are window and label artifacts, and against the first raw repricing the filter's lead is −2 min. On the nowcast gate it merely ties a Hawkes self-excitation baseline — so the latent state adds no forward-predictive value over self-excitation. — wf_mmpp_v0

This is the disciplined boundary of the whole project: the MMPP is a faithful descriptive and detection model of repricing intensity, and the amortized VI is a faithful inference method — but neither is a predictive claim. The honest characterisation, an absorption-aware intensity model whose latent state is real, detectable, and bounded by an identification ceiling, is the thesis's result whichever way the central hypothesis fell.

8 · Methodology

Rigor, nulls, and self-correction

The methodological spine is the part a technical reader should weigh most heavily. Every probe is pre-registered with named gates and a committed verdict; inference is checked against simulation; confidence intervals are market-clustered bootstrap; and findings were retracted when the evidence turned. The project's credibility rests on the negative results as much as the positive ones.

Self-corrections carried into the thesis

False positives caught and discarded

Results summary

ProbeQuestionKey numberVerdict
EfficiencyDoes the price beat a base-rate prior?Brier 0.0585 vs 0.166efficient
Directional edgeAny out-of-sample directional edge?0 / 192 lead-lagno edge
Variance regimeDo volatility regimes persist?0 / 10 marketsfalsified
Intensity clusteringIs repricing non-Poisson and clustered?10 / 10, CV 2.2–5.3established
MMPP (M1)Latent hot/cold state with dwell?beats M0 10/10established
Hierarchical poolDoes partial pooling help?+11,400 nats, 34/34go
Freeze (M2)Does intensity change as τ→0?rate-ratio 0.15 / spike 5.77×two regimes
Latent trigger axisDo trigger types separate?SNR 0.22, p=0.66bounded null
Walk-forwardDoes the latent state forecast?ties Hawkesdetect ≠ anticipate

All figures and numbers on this page are drawn from committed results in the project repository; the originating probe is named beside each. The companion chapter framework and slide deck give the higher-level views.