MSc thesis · UCL · in write-up

Latent Dynamics of Prediction-Market Price Paths

What governs how, and how fast, a prediction market reprices — and what happened when I tested my own answers hard enough to break several of them.

This is the long version, written for a reader who already wants the detail. The research page has the plain-language summary and the figures; the glossary defines every term used here. In one sentence: the question is whether a market has a hidden mood that turns busy before its price moves, or whether price moves simply cause more price moves, and the answer I found is the second.

Supervised by Prof. Philip Treleaven, University College London.
This page describes an unsubmitted MSc thesis, in write-up toward September 2026. Nothing here has been examined or graded. The research repository is private.

The question

A price with a known ending

A prediction-market price is a terminally-absorbed process — a quantity that runs until it hits a value it can never leave. It is driven to an absorbing $0/$1 outcome at an approximately known deadline, which means every path has a knowable ending — a property almost no financial time series has. The mathematics for questions of this shape is the point process: it models when events occur rather than how large they are, and its simplest form, the Poisson process, assumes events arrive independently at a steady rate. What governs the rate at which such a market reprices is a hidden state that no observable field records, so the thesis treats it as a latent-structure learning problem: where a conventional approach would assume a number of regimes or a market type, the models here infer it.

The natural conjecture is that repricing sharpens into the deadline. The thesis's answer is that it does the opposite — intensity freezes, beneath a much larger and faster clustering process — with sports the one category that spikes into its deadline instead.

Process

How the work actually moved

Nine steps. Each one is here because of what the previous one ruled out.

The figures for these results, each with its source numbers as a downloadable CSV, are on the research page: the falling hazard and the rejected Weibull, the five lead definitions, the per-market pooling gains, and the two-state separation.

1 · Build an instrument that can see the object

Existing prediction-market data is either resolved-market snapshots or polled tapes at roughly one observation per market every 15 to 60 minutes. Repricing happens in bursts of seconds. At a 30-minute cadence the object of study — which events arrive, and when — is simply invisible. So the first contribution had to be the instrument: continuous capture of the 250 most active markets at millisecond venue timestamps, about 12,313 events per market per hour, roughly a 4,900× increase in cadence.

What it forced — reliability became an auditing problem rather than an uptime problem. Feed gaps are logged and masked, never interpolated: 19 gaps over one second, totalling 74.6 minutes across an 18-day window. One whole collection day was lost to a disk-exhaustion incident and is masked by the pre-registered gap machinery as a single 45.3-hour gap, rather than by a rule written after the fact. Own downgrade — the 0.9992 median touch-match reconstruction check is "an internal-consistency check on the replay procedure rather than a validation against the trade tape."

2 · Ask whether the venue is beatable — five ways — and fail

Across a 1,348-market census, five independent pre-registered strategies (longshot decay, jump reversion, momentum, lead–lag, relative value) each show real gross anomalies, and every one dies net of costs under Benjamini–Hochberg correction. Of 192 out-of-sample lead–lag tests, not one survives. An arbitrage scan over roughly 730M book events found 12 gross partition arbitrages; one survived fees, for 36 milliseconds.

Scope, in the thesis's own words — "It covers the strategies tested, on the taker side, over one collection window; it is a statement about what was tried, not a proof that nothing works." What it forced — the turn from prediction to price formation. The efficiency chapter is demoted to the motivation control rather than a contribution.

3 · Attack my own calibration number, and watch a test collapse

The venue's late price is well calibrated — but a price minutes from resolution can be near-mechanically correct, so the interesting question is calibration at real forecast horizons. A separate pre-registered run tested exactly that, and its own cohort failed: the historical corpus has a median recorded market life of about ten minutes, so zero markets survive at 7 days or 24 hours. The first finding was negative and concerned the corpus rather than the venue.

What it forced — a disclosed cohort deviation onto the high-activity sample, run first at the pre-registered n=105 and then at n=318 under the identical frozen rule. The book mid then beats the base rate at every horizon from 24 h to 1 h, every CI (credible interval — the range the true value plausibly sits in) excluding zero. Reported against interest — the margin got smaller on the larger corpus because composition eased the benchmark, and the thesis instructs readers to quote the smaller figure.

4 · Put the obvious model up first and let it fail

The natural econometric description is a Gaussian regime-switching model of return variance. It won the wrong contest and lost the right one: it beat a random walk on held-out density on 10/10 markets, but median regime dwell was about one bin at 1 s, 10 s and 60 s alike — the dwell criterion passed 0/10. A purpose-built rescue, a sticky prior pushed to κ = 10⁴, still flickered on 9 of 10 markets with held-out likelihood essentially flat.

Disowned in the same sentence it was reported — the apparent win "is a fat-tail emission win, because a two-component variance mixture fits heavy-tailed returns better than a single Gaussian regardless of any dynamics." What it forced — relocating the latent structure from the variance of returns to the rate of repricing, and replacing the event instrument itself after the committed threshold was shown to have zero power on coarse-tick markets.

5 · Test the absorption hypothesis — and get the opposite

The pre-registered absorption test gated itself on recovery and identifiability before any empirical number was read, with both directions and the null admissible. Intensity freezes: β_τ = −0.131, CI [−0.188, −0.074], 8 of 9 markets frozen, pooled rate-ratio 0.15 CI [0.02, 0.59]. Re-anchoring at the settlement record makes the freeze stronger: 9 of 9, pooled 0.052.

What it forced — a bounded price near 0 or 1 mechanically has less room to move, so a mechanical room-to-move control had to be built as an explicit circularity guard, and the guard's own validity had to be recovery-tested before use.

6 · The primary test fails, and the guard turns out near-collinear

The primary pre-registered comparison — consensus beyond calendar — returned −1.1 nats, CI [−394.6, +409.1] at n=318. The guard passed (+6,731.0 nats, CI [+3,889.8, +10,343.2]), but the two covariates correlate at r = −0.998, and both single-driver contrasts are null.

The committed results package for this study had recorded a much stronger headline. The thesis overrides its own committed report: the guard "adjudicates a functional-form difference… It does not adjudicate 'belief consensus' versus 'mechanics' as two independent economic forces", and the surviving margin is carried by the calendar term the control omits, not by consensus in isolation.

Deliberate omission — depth data was available and was not used, because "a depth-based test would restate the identification problem rather than resolve it."

Per-episode detection lead Distribution of lead in minutes across 48 held-out hot episodes. Against the first raw repricing the median lead is -2.0 minutes. lead -11 min: 1 episode(s) -11 lead -10 min: 1 episode(s) lead -9 min: 1 episode(s) lead -7 min: 1 episode(s) -7 lead -5 min: 1 episode(s) lead -3 min: 1 episode(s) lead -2 min: 1 episode(s) -2 lead 0 min: 1 episode(s) lead 1 min: 4 episode(s) lead 2 min: 6 episode(s) 2 lead 3 min: 9 episode(s) 3 lead 4 min: 1 episode(s) lead 7 min: 5 episode(s) 7 lead 8 min: 1 episode(s) lead 9 min: 4 episode(s) lead 10 min: 4 episode(s) 10 lead 11 min: 6 episode(s) minutes early (negative = the model was late) 9 episodes

Chart scrolls sideways on a narrow screen — or open “Show the numbers” below.

Show the numbers
episode indexlead min vs causal detector
12.0
23.0
32.0
411.0
52.0
62.0
73.0
87.0
97.0
102.0
1110.0
127.0

First 12 of 48 rows. Download all 48.

Where the detections actually land. Lead in minutes for each held-out hot episode, measured against the causal detector. The mass sits close to zero and the distribution is wide; measured instead against the first raw repricing the median is −2 minutes and not one of the 48 episodes is early. n = 48 episodes · data 2026-06-17 · pre-registered · held-out, walk-forward · lead-episodes.csv, lead-metrics.csv

7 · Adjudicate the mechanism, then test the flank I left open

A pre-registered comparison found a self-exciting model beating the two-state latent-state model by +48,970 nats, CI [+35,238, +67,016], zero drop-one flips, robust across 30/60/120 s bins. Then, after the pre-registration closed, I tested the flank I had disclosed: a three-state competitor reverses the ordering by 37,868 nats. A simulation control (12/12 under a self-exciting truth) rules out a parameter-count artifact.

Claim narrowed, not defended — the scope became "the tested two-state regime model rather than the regime class". Mechanism language held back — the licensed wording is "consistent with self-excitation", never identified endogeneity, because a self-exciting kernel and serially-correlated exogenous information generate the same counts.

8 · Reconcile the two findings by test, not by argument

A further pre-registered test looked inside the winning model's own residuals for the freeze. It is there: median γ_c = −0.0391, CI [−0.0528, −0.0273] at the pre-registered n=105, and −0.0401, CI [−0.0512, −0.0314] on the final corpus. The predictive leg was null at n=105 and resolves at n=318 to +539.5 nats — about 1.1% of the self-excitation margin.

So: two scales. Self-excitation carries essentially all the short-run predictive content; the freeze is a real but slow, venue-wide quieting that survives underneath it as a measurable sliver.

Amendment disclosed — the specified gating statistic failed simulated false-positive control at 35% against a 5% bar, so an already-registered alternative was promoted before any empirical statistic was computed, under its own commit hash. Headline dropped on purpose — a larger simulation-amplified version of this effect was discarded as too fragile to translate. The headline is the measured coefficient, not the amplified one.

9 · Put my own model class on trial, and lose

A pre-registered neural benchmark, falsification-tested in both directions first and with the test set evaluated exactly once, beat the interpretable champion: +15,600 nats, CI [+10,675, +21,640], ahead on 51/51 test markets. The interpretable ladder still covers 90.8% of the total gap with 3 to 6 parameters.

The residual was localised rather than waved away: a model restricted to the champion's own counts-only information set still wins by +7,756 nats, which indicts the shape of the excitation kernel specifically — a concrete agenda rather than a black-box advantage.

Reported in the direction that hurts — the learning curve is still rising, so +15,600 is a lower bound on the neural advantage and 90.8% a upper bound on the interpretable share.

Results

The figures, with their status

Every figure below carries the protocol that produced it. Where a number exists at two corpus sizes, the pre-registered value and the final value are both given, because they differ.

Headline results and their evidential status. CIs are market-clustered.
Result Figure Status
Freeze toward end of activity β_τ = −0.131 [−0.188, −0.074] Pre-registered, gated, stable n=9 → 318
Primary test: consensus beyond calendar −1.1 [−394.6, +409.1] Pre-registered — failed, reported as failure
Circularity guard +6,731 [+3,890, +10,343] Passed, but covariates collinear at r = −0.998
Self-exciting vs two-state model +48,970 [+35,238, +67,016] Pre-registered, scoped to the tested two-state model
Three-state competitor +37,868 [+30,156, +46,662] Exploratory, post-freeze, not pre-registered
Freeze in the winner's residuals γ_c = −0.0401 [−0.0512, −0.0314] Pre-registered, one disclosed pre-outcome amendment
Neural benchmark residual +15,600 [+10,675, +21,640] Pre-registered; single frozen split, n=206 — a lower bound
Interpretable share of the gap 90.8% [88.2, 92.7] Same split; an upper bound, not intrinsic
Sports deadline-spike β_τ = +0.258 [+0.161, +0.357] Pre-registered split, Benjamini–Hochberg-corrected across categories (a correction for testing many categories at once)
Latent trigger axis SNR 0.22, permutation p = 0.66 Pre-registered — null, "reported without salvage"

Limitations

Limits of the evidence

Each limitation below is drawn from the thesis's own caveats. The committed CAVEATS documents they come from are written but embargoed until submission; the method section says what that covers.

  • The guard cannot separate belief from mechanics. At r = −0.998 the two covariates are near-identical transforms of the same price. The comparison distinguishes two functional forms of price extremity, and nothing stronger.
  • The model-selection result is not stable under flexibility. Grant the latent-state side a third state and the held-out ordering reverses. One-step-ahead count likelihood does not separate the two accounts.
  • Every model in the ladder fails absolute fit. A time-rescaling test rejects the winner on most markets. It is the best-performing tested approximation, not a generatively adequate description.
  • The branching-ratio estimate is biased. Estimates of this kind are biased toward one when the baseline is nonstationary and the kernel is misspecified — and this panel is arguably both, by the thesis's own evidence. It is a point estimate subject to a known upward bias, not a claim that these markets sit at criticality.
  • The cross-venue comparison is n=1 and the second venue was on a delayed feed, so no lead or lag magnitude is claimed. The delay-independent part is structural: one venue repriced goals in under a second while the other suspended trading for one to three minutes, with no tradeable basis between them.
  • The corpus is narrow. Six weeks, one venue, its liquid actively-repriced head — 88% of eligible volume but only 42% of eligible markets — over an event mix that included a World Cup. Cohorts are nested, not independent, and the thesis says so.
  • Reproducibility is partial. The repository is private. No library versions, containers or lockfiles are pinned. Not every figure comes from a pre-registered component.

← Back to research