Skip to main content

Possible MPhil/PhD Projects in Statistics

Suggested projects for postgraduate research in statistics.

Areas of expertise

Below is an indicative list of research projects which our staff would be interested in supervising. The list is not exhaustive, and we encourage interested applicants to also check the webpages of individual academics. If you are interested in any of these projects, or would like to discuss a project with someone, please feel free to approach the prospective supervisor by email. Please also check the details of our research degrees. We welcome applications online.

 

For further information, please contact the PG tutor/selector in Statistics Axel Fincke.

Assurance in reliability demonstration testing

In a reliability demonstration test the producer of a hardware product demonstrates to a consumer that the product meets a certain level of reliability. As most hardware products have very high reliability, such tests can be prohibitively expensive, requiring large sample sizes and long testing periods. Accelerated testing can reduce the testing time, but introduces the additional complication of having to infer the relationship between failure times of the stressor variable at accelerated and normal operator conditions.

Previous attempts to plan and analyse reliability demonstration tests have utilised power calculations and hypothesis tests or risk criteria. More recently, Wilson & Farrow (2019) proposed the use of assurance to design reliability demonstration tests and suitable Bayesian analyses of the test data. Assurance provides to unconditional probability that the reliability demonstration test will be passed. Work to date has focussed on Binomial and Weibull observations. This project would extend the use of assurance to design reliability demonstration tests, considering a wider class of failure time distributions and implementing an augmented MCMC scheme to evaluate the assurance more efficiently.

Supervisor: Kevin Wilson

Bayesian design and analysis of clinical trials using assurance

The standard way to choose a sample size a clinical trial is via a power calculation – we choose the smallest sample size that gives us 80% or 90% power to detect a certain treatment effect, based on a hypothesis test on the treatment effect at the end of the trial. This requires various model parameters to be known a priori, particularly for more complex trial designs such as cluster randomised trials. If instead we use prior distributions on the model parameters, then we can integrate over our uncertainty on them and calculate the Bayesian power, known as assurance, for the sample size calculation. In this case it seems sensible to also conduct a Bayesian analysis following the trial. Such a fully Bayesian design and analysis has been proposed for two arm cluster trials with a continuous outcome. This project would extend the approach to more complex trial designs such as crossover trials, stepped wedge trials and adaptive trials.

Supervisor: Kevin Wilson

Dynamic Bayesian modelling of endurance sports

For runners in long-distance races such as marathons, accurately predicting the finishing time is crucial for selecting the pacing strategy which leads to the best possible performance: Run too fast and you will “hit the wall”. Run too slow and you won’t reach your full potential. Similar considerations apply in other endurance sports, e.g., cycling or rowing, or even in non-sports contexts such as in the military.

 

Conventional approaches for finishing-time prediction are static, i.e., they only make predictions before the race. Once the race has started, they can neither account for an over-optimistically chosen initial pace nor adapt to unexpected factors affecting performance. A small number of dynamic approaches exist in the literature but these are often based on mathematical models which are too simplistic.

 

This project aims to predict the finish time combined with a systematic quantification of uncertainty and to update the prediction in real time as new data come in. To that end, you will develop dynamic Bayesian models which make use of some of the data collected by fitness trackers/GPS watches during a race, e.g. pace, heart-rate or elevation data, along with other covariates. You will also develop suitable computational statistical methods (e.g. based around sequential Monte Carlo methods) which can be used to update the prediction as new data become available throughout the race.

Supervisor: Axel Finke

Models for bidispersed count data

Abstract: Classic models for count data can readily accommodate overdispersion relative to the Poisson model, but models for underdispersed counts - where the mean exceeds the variance - are less well established, and those that have been proposed are often hampered by, for instance: lack of natural interpretation, a restricted parameter space, computational difficulties in implementation or some combination of all three. At the individual level, one can often encounter both over- and underdispersion, or bidispersion, within the same dataset and failure to allow for this bidispersion leads to inferences on parameters that are either conservative or anti-conservative. In this project, models to handle such bidispersed data, typically at an individual level, will be developed and applied to case studies drawn from sport and the social sciences.

Supervisor: Pete PhilipsonDaniel Henderson

Post-Bayesian Statistical Methods

Mathematics is a powerful tool to explain, to reason about, and to some extent predict, complex phenomena in the real world. As part of this process, we often want to tune the parameters of these models to most accurately match an observed dataset. A reflexive strategy is to adopt the Bayesian approach, which enables probabilistic uncertainty quantification in subsequent predictions based on the mathematical model. Yet, mathematical models are designed to prioritise tractability and insight, not parsimony with respect to an observed dataset. This has major implications; given enough data, predictions can be simultaneously high-confidence and incorrect. This project will unpack the dilemma and explore possible resolutions based on recent advances in post-Bayesian methodology.

Supervisor: Professor Chris Oates

Seamless and adaptive designs in diagnostic test development

Approximately 70% of clinical decisions are influenced by the use of in vitro diagnostics (IVDs), and conception to adoption of new diagnostics takes approximately 10 years. The ability to speed up the development of new diagnostics, including IVDs, is critical for the long-term sustainability of the UK diagnostics sector and efficiency of the NHS. This project would aim to reduce the time to market of diagnostics by developing novel Bayesian methods for the design and analysis of diagnostic studies, with a particular focus on post-marketing utility studies. Through development of seamless and adaptive designs for diagnostic studies, we can make best use of data to save time and money in test development and enable seamless transitions between stages. Adaptive designs will allow efficacy or futility of diagnostics to be determined more quickly. The project will apply the developed methods to assess real diagnostic devices that have been recently approved.

Supervisor: Kevin Wilson

Spatio-temporal modelling of house price index with open data

House price indices reflect the average appreciation of houses of an area at a time point. Traditionally, they can only be calculated with the information about the individual characteristics of houses, which is usually withheld by private banks and mortgage lenders. On the other hand, the repeat sales regression model, developed in econometrics, can be used even in the absence of such information. An example data set of such nature is the house price data from Land Registry, to which we will apply a hierarchical statistical model to compute the indices. 

 

Within the model, spatial dependence will be incorporated to account for Tobler's first law that near things are more related to each other, while a temporal component will account for the seasonality and autoregressive nature of house price indices, which will be treated as unobserved variables. When it comes to statistical inference, efficient computational Bayesian methods such as integrated nested Laplace approximation (INLA) or Markov chain Monte Carlo (MCMC) will be used.

Supervisor: Dr Clement Lee

Statistical methods for precision medicine

Heterogeneity among patients commonly exists in clinical studies, presenting challenges in medical research. It is widely acknowledged that the population comprises diverse subtypes, each distinct from the other. Precision medicine research aims at identifying the sub-types and thus tailoring disease prevention and treatment. The primary research challenge is to identify the sub-population types and determining the prediction variables for each sub-type. This PhD research project will explore how to integrate classification methods and variable selection techniques to effectively address these complexities.

Supervisor: Prof Hongsheng Dai

Stochastic processes in spaces of evolutionary trees

Phylogenetic or evolutionary trees are inferred from genetic sequence data and are a vital ingredient of many applications in molecular biology. During the inference process samples of trees are often obtained, but characterising and analyzing such samples is challenging because the space of all trees with a fixed leaf-set is not a Euclidean vector space. While various notions of phylogenetic tree space exist, there is often a lack of probabilistic machinery in these space. The project supervisor has developed a family of distributions in the well-known Billera-Holmes-Vogtmann (BHV) tree space which are akin to Guassian distributions, via kernels of Brownian motion, together with methods for fitting these distributions to samples of trees. The aim of this project is to further exploit this idea by developing a wider range of statistical models in BHV tree space with Gaussian kernels as the noise model, or by developing analogous methods in wald space. Wald space is an alternative to BHV tree space with a more complex geometry and for which statistical models are in their infancy. This project will involve development of the theory of stochastic processes and computational methods for performing inference in non-Euclidean geometries.

Supervisor: Tom Nye

Find out about the statistics research group