提出Anchor PCA,让多领域数据共享更稳健的低维表示。
Anchor PCA
- 通过优化共享与各领域特有嵌入的一致性,寻找共变方向。
- 在模拟和真实气体传感器数据上,对未见领域解释方差更高。
- 适合处理存在领域偏移的多源数据,如时序漂移场景。
主成分分析(PCA)是广泛应用的无监督降维方法。本文研究多相关领域数据的PCA问题。由于各领域主成分通常不同,传统方法将数据合并后做PCA以获得共享低秩嵌入,但可能聚焦于仅在少数领域中变化大的虚假方向。为获得在未见相似领域仍能解释大部分方差的鲁棒嵌入,我们提出锚定主成分分析(Anchor PCA),其核心是权衡总体解释方差与共享嵌入和领域特异嵌入之间的一致性。该方法等价于对一个修正目标矩阵进行PCA,可高效求解。此外,我们证明Anchor PCA可恢复最大不变子空间,并在领域特异性协方差受界膨胀条件下具有极小极大重构意义。在模拟数据及存在时序漂移的真实气体传感器数据上,分别验证了其能恢复最大不变子空间,且在未见领域解释方差高于数据池化基线和最坏情况替代方案。这些结果确立了Anchor PCA作为多域数据鲁棒无监督降维的有力方法。
原文摘要 · Abstract (English)
Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques. We study PCA for data from multiple related domains. Since principal components generally differ across domains, one way to obtain a shared low-rank embedding is to perform PCA on the pooled data. However, this approach can focus on spurious directions that exhibit high variation in only a few domains. To find a robust embedding that still explains most variance in unseen but similar domains, we propose instead to focus on shared directions of variation. To this end, we introduce Anchor PCA which trades off overall explained variance with agreement between the shared and domain-specific low-rank embeddings. Anchor PCA amounts to PCA on a modified target matrix and thus can be solved efficiently. Moreover, we show that Anchor PCA recovers a maximal invariant subspace and admits a minimax reconstruction interpretation under bounded domain-specific covariance inflations. On simulated and real-world gas sensor data with temporal drift, we demonstrate, respectively, that Anchor PCA recovers the maximally invariant subspace and yields embeddings that explain more variance on unseen domains than the pooling baseline and a worst-case alternative. Taken together, these findings establish Anchor PCA as a promising approach to robust unsupervised dimension reduction from multi-domain data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。