提出稳定PCA,从多源数据中学习鲁棒低维表示
StablePCA: Distributionally Robust Learning of Shared Representations from Multi-Source Data
- 通过最大化最差情形解释方差,构建跨源稳定的低维表示
- 设计凸松弛与镜面逼近算法,实现全局收敛求解
- 提供数据依赖的验证证书,适用于存在系统偏差的数据
在融合多源高维数据时,核心目标是提取能在不同数据源间有效近似原始特征的低维表示,以发现可迁移结构并缓解批次效应等系统偏差。本文提出分布鲁棒的稳定主成分分析(StablePCA),通过最大化多个数据源上的最差情形解释方差来构建稳定的潜在表示。经典PCA扩展到多源场景的主要挑战在于非凸的秩约束,导致StablePCA为非凸优化问题。为此,我们对StablePCA进行凸松弛,并开发一种高效的镜面逼近算法求解松弛问题,具备全局收敛性保证。由于松弛问题通常不同于原问题,我们进一步引入数据依赖的验证证书,用于评估算法对原非凸问题的求解质量,并给出松弛紧致的条件。最后,我们探讨了基于不同损失函数的多源PCA分布鲁棒变体。
原文摘要 · Abstract (English)
When synthesizing multi-source high-dimensional data, a key objective is to extract low-dimensional representations that effectively approximate the original features across different sources. Such representations facilitate the discovery of transferable structures and help mitigate systematic biases such as batch effects. We introduce Stable Principal Component Analysis (StablePCA), a distributionally robust framework for constructing stable latent representations by maximizing the worst-case explained variance over multiple sources. A primary challenge in extending classical PCA to the multi-source setting lies in the nonconvex rank constraint, which renders the StablePCA formulation a nonconvex optimization problem. To overcome this challenge, we conduct a convex relaxation of StablePCA and develop an efficient Mirror-Prox algorithm to solve the relaxed problem, with global convergence guarantees. Since the relaxed problem generally differs from the original formulation, we further introduce a data-dependent certificate to assess how well the algorithm solves the original nonconvex problem and establish the condition under which the relaxation is tight. Finally, we explore alternative distributionally robust formulations of multi-source PCA based on different loss functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。