Séminaire Images Optimisation et Probabilités
Jérémie Bigot
( IMB U-Bordeaux )Salle de Conférences
19 novembre 2026 à 11:15
Principal Component Analysis (PCA) is one of the fundamental tools for statistical dimension reduction. Nevertheless, no canonical extension of PCA to distribution-valued data has emerged within the framework of optimal transport. While Wasserstein barycenters and Wasserstein least-squares regression have recently been shown to arise as canonical liftings in the Wasserstein space of Euclidean averaging and standard linear regression, an analogous lifting principle for PCA has not been considered so far. In this paper, we introduce Lifted PCA, a canonical extension of Euclidean PCA to the Wasserstein space obtained through the framework of lifting functionals in the space of probability measures by convex extension. Starting from the classical variational formulations of PCA on the Stiefel and Grassmann manifolds, we derive two equivalent minimization problems over probability measures supported on these manifolds that correspond to Wasserstein barycenters with generalized costs. We prove that these formulations are linked by an exact analogue of the equivalence between the Stiefel and Grassmann formulations of Euclidean PCA, thereby preserving the original variational structure of PCA on these manifolds after lifting. Building upon the generalized Wasserstein barycenter problem, our analysis further shows that Lifted PCA admits
dual formulations and optimality conditions for its characterization, as well as equivalent multimarginal optimal transport and matching-for-teams formulations. We also establish statistical consistency for empirical estimators and develop computational algorithms based on the Wasserstein barycenter formulation over the Stiefel manifold. Our results show that while Euclidean PCA summarizes a collection of vectors by projections onto an optimal low-dimensional linear subspace, Lifted PCA summarizes a collection of probability distributions using an optimal probability measure supported on multiple low-dimensional subspaces. Numerical experiments on distributional datasets illustrate the practical performance of the proposed methodology and its benefits for distributional data analysis.