MSc Machine Learning · Thesis
Learning latent repricing dynamics in terminally-absorbed price processes — non-parametric Bayesian and variational point-process models, with high-frequency Polymarket order-book data as the testbed.
Juan Mediavilla · University College London · MSc Machine Learning
Academic Supervisor: Prof. Philip Treleaven
Technical Advisor: TBC
Abstract
This thesis investigates original research in the machine learning of latent dynamics in terminally-absorbed price processes: the high-frequency price paths of prediction markets, which are driven to an absorbing $0/$1 outcome at an approximately known deadline. What governs how — and how fast — such a market reprices is a hidden state that cannot be hand-specified; it must be learned from data. The thesis therefore treats price-path dynamics as a latent-structure learning problem rather than a parametric estimation problem.
Price-formation dynamics have traditionally been modelled in econometrics, through regime-switching and stochastic-volatility specifications that fix a parametric form in advance. This thesis instead investigates the use of machine-learning algorithms — unsupervised latent-variable models, non-parametric Bayesian mixtures, and variational inference — to learn the generative structure of repricing directly. These methods are adopted for their non-parametric nature: they require no explicit assumption on the data-generating process of a venue whose paths are bounded, terminally absorbed, and shaped by exogenous information arrival. Where a regime, a cluster count, or a participant type would otherwise have to be assumed, it is instead inferred.
The thesis comprises three method-led investigations. Experiment 1 — supervised baselines & calibration fixes the framing: directional classifiers and Brier-score calibration of coarse priors against the market establish that the venue is near-efficient gross of fees and that no prior beats the market price — the negative result that motivates the unsupervised turn from prediction to price formation. Experiment 2 — unsupervised latent-state model selection decisively falsifies a natural variance-regime hypothesis (apparent regimes flicker at every timescale) and instead recovers, with an unsupervised Markov-modulated Poisson process, a hidden hot/cold intensity state with genuine dwell. Experiment 3 — the absorption-aware latent-intensity model extends that process into one whose latent intensity is modulated by time-to-resolution — a model class with no standard analogue — fitted by variational inference and compared against a self-exciting (Hawkes) generative alternative; a Dirichlet-process mixture, inferred by amortized variational inference, learns the number and shape of latent market and participant types without assuming it, over a hierarchical Bayesian backbone pooled by Metropolis-within-Gibbs sampling, with variational and MCMC inference cross-validated against each other.
Prediction markets serve as the testbed: a purpose-built, gap-audited, millisecond-resolution corpus of the 250 most active Polymarket markets, captured continuously through resolution, supplies the path-to-absorption ground truth against which the learned models are validated. The work is designed to be defensible whatever the empirical outcome — a rigorous null is reported as readily as a positive one. The dissertation's original contributions to science are: an absorption-aware latent-intensity model class for terminally-absorbed processes; the use of amortized variational inference and non-parametric Bayesian mixtures to learn market and participant structure that cannot be hand-labelled; and a reusable high-frequency dataset and pipeline. No trading capital is deployed; the venue is independently near-efficient gross of fees, which is precisely what makes price formation, not prediction, the object of study.
Contributions
A new latent-variable model class for terminally-absorbed processes: a Markov-modulated Poisson process whose hidden intensity is modulated by time-to-resolution, fitted by variational inference — no standard analogue in financial time-series.
Unsupervised learning of latent market and participant structure with Dirichlet-process mixtures and amortized variational inference — the cluster count is inferred, not assumed — over a hierarchical Bayesian backbone.
A gap-audited, millisecond-resolution order-book corpus of Polymarket, and the 24/7 collection pipeline that produces it — the testbed on which the learned models are validated.
Read on
The full abstract and eight chapter micro-abstracts — context and data, the empirical discovery, the model the discovery dictates, then validation — with datasets and risks.
Open framework →A figure-led deep-dive for technical readers: the data, the efficiency findings, the learned latent-state model (MMPP + variational inference), the non-parametric Bayesian latent-types flagship, and the inference machinery — with the real numbers.
Open technical →A navigable ten-slide web deck: motivation, background, objectives, methodology, the three experiments, contributions and impact — with real figures and results.
Open deck →