Events

Upcoming

Formal Bayesian Transfer Learning via the Total Risk Prior

Wednesday, September 9, 2026

Speaker: Dr. Nathan Wycoff (University of Massachusetts Amherst)

Existing methods for transfer learning struggle to deal with situations where the source datasets are limited and not guaranteed to be well-aligned with the target dataset. A typical strategy is to use the empirical loss minimizer on the source data as a prior mean for the target parameters. Our key conceptual contribution is to use a risk minimizer conditional on source parameters instead. This allows us to construct a single joint prior distribution for all parameters from the source datasets as well as the target dataset. As a consequence, we benefit from full Bayesian uncertainty quantification and can perform model averaging via Gibbs sampling over indicator variables governing the inclusion of each source dataset. We show how a particular instantiation of our prior leads to a Bayesian Lasso in a transformed coordinate system and discuss computational techniques to scale our approach to moderately sized datasets. We discuss connections between the Maximum a Posteriori estimate associated with our approach and the recently proposed Trans-Lasso method and demonstrate that the MAP estimator MSE-dominates the Trans-Lasso in the normal means setting when there is no regularization on the source datasets. Finally, we perform numerical experiments finding that full Bayesian inference provides superior predictive performance relative to Trans-Lasso on a genetics application, especially when the source data are limited.

Past

Probabilistic approaches for fair clustering

Monday, June 22, 2026

Speaker: Dr. Abhisek Chakraborty (Eli Lilly and Company)

The advent of ML-driven decision-making has led to an increasing focus on algorithmic fairness. The widespread utility of clustering has naturally prompted the proliferation of literature on fair clustering. A popular notion of fairness in clustering mandates the clusters to be balanced, i.e., each level of a protected attribute must be approximately equally represented in each cluster. In this talk, I offer a novel model-based formulation of fair clustering, complementing the existing literature which is almost exclusively based on optimizing appropriate objective functions. We first rigorously define a notion of fair clustering at the population level and develop a Bayesian methodology equipped with a novel hierarchical prior specification that targets the population level object by enforcing the notion of balance in the resulting clusters. In addition, we devise a scheme for principled performance evaluation of competing algorithms leveraging on a concrete notion of optimal recovery. An efficient collapsed Gibbs sampler is developed to sample from the posterior by integrating a novel scheme for non-uniform sampling from the space of binary matrices with fixed margin with a proposal guided by optimal transport. Superior empirical performance of the proposed methodology, compared to the state-of-the-art, is demonstrated across numerical experiments, benchmark data-sets, and gender-neutral fair clustering in the distress analysis interview corpus.

Applications of Bayesian modelling in studies of climate, health, and equity

Wednesday, March 25, 2026

Speaker: Dr. Robbie M. Parks (Columbia University)

A major component of my research focuses on quantifying the health impacts of climate-related hazards and modelling population dynamics using detailed and large datasets, on scales ranging from small-area to multi-country, for which Bayesian modelling can afford numerous advantages. In this seminar, I will highlight some of my major research efforts on climate, health, and equity, including several recent and ongoing studies in the United States, Chile, and the Philippines. Focus topics will include natality and mortality disparities, studies of the association of health-relevant outcomes with heat stress, tropical cyclones, and wildfires, and ongoing work on climate change and health attribution.

A Bayesian Analysis of Spike Activity in Organoids Under Electrical Stimulation

Wednesday, March 11, 2026

Speaker: Dr. Babak Moghadas (Johns Hopkins University)

Characterizing how neural organoids respond to electrical stimulation is essential for understanding their computational and adaptive properties. In this study, we develop a Bayesian hierarchical framework to analyze spike activity recorded from organoids using multi-electrode array (MEA) platforms under varying stimulation intensities. Spike counts are modeled as stochastic processes, allowing us to quantify stimulation effects while accounting for electrode-level variability, repeated trials, and parameter uncertainty.

Using Bayesian generalized linear modeling and posterior inference, we estimate stimulation-dependent changes in firing rates and obtain uncertainty quantification. The hierarchical structure enables partial pooling stimulation levels, improving statistical efficiency and robustness.

Our results show systematic modulation of spike activity with increasing stimulation intensity, alongside substantial heterogeneity across recording sites. This probabilistic framework provides a principled and flexible approach for analyzing organoid electrophysiology data and supports inference in novel biological neural systems. 

Bayesian Transfer Learning Approaches for Large-scale Spatiotemporal Problems

Wednesday, February 18, 2026

Speaker: Luca Presicce (University of Milan, Bicocca)

The increasing availability of large-scale geospatial and spatiotemporal data presents new opportunities and challenges for statistical modeling in environmental, technological, medical, and other complex areas, which increasingly rely on massive multivariate spatiotemporal datasets. Yet, Bayesian learning for such problems remains severely limited by computational bottlenecks and the lack of flexible modeling tools. Modern applications require methods that are adaptive and effective, but still computationally efficient, scalable to massive datasets, and capable of delivering reliable automated inference with principled uncertainty quantification and (possibly) minimal experienced human intervention. Classical Bayesian approaches, although theoretically appealing and offering rich inferential frameworks, often become computationally infeasible in data-rich environments, especially when confronted with massive datasets or dynamic, high-dimensional dependence structures. Existing approaches often fail to scale, leaving a gap between the theoretical richness of Bayesian inference and its practical deployment in data-rich applications. This thesis develops Bayesian transfer learning methodologies to address these challenges, enabling efficient information propagation and scalable inference across large spatial and spatiotemporal domains, providing a unified framework that merges distributional theory for matrix-variate models with computational innovations in Bayesian predictive stacking. Through extensive simulation experiments and data applications to global and satellite monitoring of vegetation indices, sea surface temperature, and land-atmospheric climate composition, the thesis also demonstrates the potential of Bayesian transfer learning to redefine spatial and spatiotemporal multivariate modeling, providing flexible, computationally efficient solutions that open the way for scalable, automated, and truly modern tools for geospatial learning in data-rich environments.