arXiv:2603.11304stat.MLcs.AI2026-03被引 1

针对跨领域数据的低秩近似,提出更鲁棒的worst-case优化方法。

Worst-case low-rank approximations

  • 构建统一框架wcPCA,优化最差情况下的性能表现
  • 在真实数据上,最差情况性能显著提升,平均性能损失小
  • 适用于医疗、环境等分布异质场景,适合追求稳健性的研究者

健康、经济与环境科学中的真实数据常来自异质领域(如医院、地区或时间段)。在此类场景中,分布偏移会导致标准PCA不可靠——例如,主成分在未见领域解释的方差远低于训练域。现有方法(如FairPCA)尝试优化多个领域的最差表现。本文提出统一框架wcPCA,拓展至其他目标,得到新型估计器如norm-minPCA和norm-maxregret,更适用于总方差异质的应用。证明所有目标下,估计器不仅在观测源域,且在所有协方差位于(可能归一化后)源协方差凸包内的目标域上均达到最差情况最优。建立了经验估计器的一致性与渐近最差情况保证。将方法扩展至矩阵补全问题,证明归纳式矩阵补全具有近似最差情况最优性。模拟及两个真实世界应用(生态系统-大气通量)显示,最差情况性能明显改善,平均性能仅轻微下降。

原文摘要 · Abstract (English)

Real-world data in health, economics, and environmental sciences are often collected across heterogeneous domains (such as hospitals, regions, or time periods). In such settings, distributional shifts can make standard PCA unreliable, in that, for example, the leading principal components may explain substantially less variance in unseen domains than in the training domains. Existing approaches (such as FairPCA) have proposed to consider worst-case (rather than average) performance across multiple domains. This work develops a unified framework, called wcPCA, applies it to other objectives (resulting in the novel estimators such as norm-minPCA and norm-maxregret, which are better suited for applications with heterogeneous total variance) and analyzes their relationship. We prove that for all objectives, the estimators are worst-case optimal not only over the observed source domains but also over all target domains whose covariance lies in the convex hull of the (possibly normalized) source covariances. We establish consistency and asymptotic worst-case guarantees of empirical estimators. We extend our methodology to matrix completion, another problem that makes use of low-rank approximations, and prove approximate worst-case optimality for inductive matrix completion. Simulations and two real-world applications on ecosystem-atmosphere fluxes demonstrate marked improvements in worst-case performance, with only minor losses in average performance.

低秩近似最差情况稳健学习矩阵补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。