Faculty/School

Faculty of Science

School of Mathematical Sciences

Topic status

We're looking for students to study this topic.

Research centre

Primary Supervisor

Dr Trung Tin Nguyen
Position
MACSYS Post-doctoral Research Fellow (Applied Statistics)
Division / Faculty
Faculty of Science

Other QUT supervisors

Professor Chris Drovandi
Position
Professor
Division / Faculty
Faculty of Science

Overview

How does a machine learning system decide how many specialists it needs to make good predictions? Mixture-of-experts (MoE) models (Jacobs et al., 1991; Jordan and Jacobs, 1994) models tackle hard problems by splitting them among many simple sub-models, or experts, the same idea now powering some of the largest AI language models, and they are valued for their flexibility in modelling complex, real-world data. Despite their success, one deceptively simple question is still hard to answer well: how many experts does a given dataset actually need? Too few and the model is too crude; too many and it overfits and wastes computation. Existing answers are scattered across the literature, with many competing methods but few systematic, scalable comparisons (Hai et al., 2026), which is exactly the gap this project tackles.

Mean-field variational inference (MFVI) is a fast way to fit these models: rather than exhaustively exploring every possible setting of the unknown parameters, it approximates the answer by assuming those unknowns are independent of one another, trading a little accuracy for a large speed-up. Recent theory shows that a quantity this method already computes, the evidence lower bound (ELBO), essentially a score for how well a model fits while penalising complexity, can be used to choose between models in a statistically principled way, in the sense of selection consistency (Zhang & Yang, 2024, Bariletto et al., 2026). However, because MFVI assumes the parameters are independent, it tends to be over-confident, it underestimates how uncertain its own estimates really are (statisticians call this underestimating the posterior uncertainty). Whether, and how, this over-confidence affects the reliability of ELBO-based model selection is not yet well understood, and is exactly what this project sets out to investigate.

This project will investigate scalable methods for Bayesian model selection based on variational predictive resampling (Battaglia et al., 2026), a recently developed posterior-sampling method that exploits the already-strong predictive distributions of mean-field variational inference to recover posterior dependence that mean-field VI misses and to produce more honest uncertainty estimates. Variational predictive resampling has so far been developed and demonstrated for fixed-dimension models (e.g. linear and logistic regression and hierarchical models); a central novelty of this project is to adapt and evaluate it for model selection in MoE models, where the number of experts is unknown. The central question is simple to state: does this new fix choose the number of experts more reliably than standard MFVI, without giving up the speed that made MFVI attractive in the first place? You will answer it using controlled simulation studies (where the true number of experts is known) on 2-3 settings and at least one real benchmark dataset, building on an existing MFVI MoE implementation provided by the supervisory team.

References:

Jacobs, R. A., Jordan, M. I., Nowlan, S. J., & Hinton, G. E. (1991). Adaptive mixtures of local experts. Neural Computation, 3(1), 79-87.

Jordan, M. I., & Jacobs, R. A. (1994). Hierarchical mixtures of experts and the EM algorithm. Neural Computation, 6(2), 181-214.

Bariletto, N., Nguyen, H., Ho, N., & Rinaldo, A. (2026). On Bayesian Softmax-Gated Mixture-of-Experts Models. arXiv preprint arXiv:2604.20551.

Battaglia, L., Cortinovis, S., Holmes, C., Frazier, D. T., & Jewson, J. (2026). Variational predictive resampling. arXiv preprint arXiv:2605.11168.

Hai, D. T., Mai, T. N., Nguyen, T., Ho, N., Nguyen, B. T., & Drovandi, C. (2026). Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency without Model Sweeps. In Proceedings of the 29th International Conference on Artificial Intelligence and Statistics.

Zhang, Y., & Yang, Y. (2024). Bayesian model selection via mean-field variational approximation. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(3), 742–770.

Research engagement

The student will engage with a focused literature review of recent Bayesian Mixture of Experts and variational predictive resampling papers, hands-on computational implementation and debugging, the design of simulation experiments, and the presentation of results.

Research activities

The student will work directly with the supervisory team, TrungTin Nguyen as primary day-to-day supervisor, with Christopher Drovandi, meeting at least weekly per the VRES Learning Contract. Building on an existing MFVI Bayesian Mixture of Experts codebase (Python/R) provided by the team, the student will complete the plan below over a 6–10 week project (≈120–150 hours), spanning VRES Research Session 1 (November–December, 2026) and Session 2 (January, 2027) either side of the Christmas break. A 6-week project finishes at the core comparison (Weeks 1–6); the optional later weeks deepen the empirical study and polish the final outputs.

- Weeks 1–2 (Foundations): Read 3–4 core papers (Zhang & Yang, 2024; Battaglia et al., 2026; Bariletto et al., 2026; Hai et al., 2026) and set up and run the existing MFVI MoE code on a simple simulated dataset to reproduce ELBO-based selection of the number of experts.

- Weeks 3–4 (Implement the method): Implement a variational predictive resampling routine on top of the existing MFVI fit and validate it on one synthetic example with a known number of experts.

- Weeks 5–6 (Core comparison, minimum deliverable): Run a focused simulation study comparing ELBO-based versus predictive-resampling-based selection across a handful of scenarios; produce summary figures and a short write-up.

- Weeks 7–8 (Scale up, if extending): Extend the comparison to 2–3 larger synthetic settings and at least one real benchmark dataset, recording both selection accuracy and runtime.

- Weeks 9–10 (Consolidate and present, if extending): Clean and document the code for reproducibility and prepare the final technical report or poster for the VRES final presentation.

Research skills

Through this project you will gain hands-on skills in: modern Bayesian computation and variational inference; implementing and debugging probabilistic models in R/Python; designing and running reproducible simulation studies; benchmarking statistical methods for both accuracy and computational cost; reading and distilling primary research literature; and communicating quantitative results through a written report and figures, a transferable toolkit valued in both academia and data-science industry roles. You will be mentored closely by the supervisory team.

Outcomes

The aims are to: (1) implement variational predictive resampling for selecting the number of experts in a Bayesian Mixture of Experts model; and (2) empirically compare it against standard ELBO-based MFVI selection. Expected outcomes are a documented, reproducible code implementation; a small simulation study (2-3 synthetic settings plus one real dataset) reporting selection accuracy and runtime; and a short technical report or poster summarising the findings. Strong results could seed a co-authored research manuscript, and the project is designed as a natural springboard into an Honours or PhD project with this group.

Skills and experience

Student Profile:

Eligible students must meet the Faculty of Science VRES requirement of a GPA of 6.0 or above. No prior experience with mixture models or variational inference is required, strong fundamentals in statistics/programming and a willingness to learn are what matter most.

Desired Skills:

  • Solid foundation in Bayesian statistics.
  • Experience with statistical and scientific computing in R or Python (knowledge of Rcpp/Cython/C++ is desirable but not essential).
  • Interest in computational statistics and machine learning methods.

Start date

2 November, 2026

End date

19 February, 2027

Location

Gardens Point campus and/or Zoom

Keywords

Contact

TrungTin Nguyen

0493903076

t600.nguyen@qut.edu.au