骆威,浙江大学,百人计划研究员,于2014年毕业于美国宾夕法尼亚州立大学,之后任职于美国Baruch College,于2018年加入浙江大学,获得国家高层次人才称号。骆威的研究方向包括充分降维和因果推断,在Annals of Statistics, Biometrika, JRSSB, JMLR等统计和机器学习国际学术期刊上发表了多篇论文。
报告摘要:The Gaussian Mixture Model (GMM) has been widely used for clustering analysis. It is commonly fitted by the maximal likelihood approach, which is computationally challenging due to the non-convex minimization, especially as the dimensionality grows. To address this issue, we propose a two-step approach by recovering the intrinsic low-dimensional structure of GMM under additional constraints on its heterogeneity; that is, there exists a low-dimensional linear transformation of the data, given which the rest of the data are normally distributed and thus redundant for clustering. Our approach first recovers the desired low-dimensional data based on Stein's Lemma and then uses the reduced data only to fit GMM. Its computational efficiency comes from both the lower dimensionality and denoising of the data. Under a sparsity assumption of the clustering pattern, our approach can be generalized in high-dimensional settings. With the aid of a novelly constructed pseudo response, it can also be embedded into a general framework of sufficient dimension reduction, which encompasses a wider class of methods beyond Stein's Lemma to recover the low-dimensional structure of GMM. These findings are illustrated in the numerical studies at the end.