arXiv:2601.15239stat.MLcs.LG2026-01被引 1

多场景主成分分析,挖掘跨不同情境数据的共享变化模式

Multi-context principal component analysis

  • 将主成分分析扩展到多情境数据,分解出跨子集情境共享的因子
  • 在癌症基因表达中发现与肺癌进展相关的变异轴,且仅在方差上显著
  • 适用于生物、语言等多情境数据,揭示传统方法忽略的深层规律

主成分分析(PCA)是捕捉数据变异因素的工具。当前,数据常在多个情境下收集(如不同疾病个体、不同细胞类型或文本中的词语)。尽管这些数据的变化因素在情境间必然存在共享,但尚无系统方法可恢复此类因素。本文提出多情境主成分分析(MCPCA),一个理论与算法框架,能将数据分解为跨子集情境共享的因子。应用于基因表达数据时,MCPCA揭示了多种癌症类型间的共享变异轴,以及一个仅在肿瘤细胞方差中显著、与肺癌进展相关的轴。应用于语言模型生成的上下文词嵌入时,MCPCA映射出关于人性讨论的数十年演变,揭示科学与虚构之间的对话进程。这些轴无法通过简单合并或单独分析各情境数据获得。MCPCA 是对传统 PCA 的合理推广,旨在理解跨情境数据背后的驱动因素。

原文摘要 · Abstract (English)

Principal component analysis (PCA) is a tool to capture factors that explain variation in data. Across domains, data are now collected across multiple contexts (for example, individuals with different diseases, cells of different types, or words across texts). While the factors explaining variation in data are undoubtedly shared across subsets of contexts, no tools currently exist to systematically recover such factors. We develop multi-context principal component analysis (MCPCA), a theoretical and algorithmic framework that decomposes data into factors shared across subsets of contexts. Applied to gene expression, MCPCA reveals axes of variation shared across subsets of cancer types and an axis whose variability in tumor cells, but not mean, is associated with lung cancer progression. Applied to contextualized word embeddings from language models, MCPCA maps stages of a debate on human nature, revealing a discussion between science and fiction over decades. These axes are not found by combining data across contexts or by restricting to individual contexts. MCPCA is a principled generalization of PCA to address the challenge of understanding factors underlying data across contexts.

主成分分析多情境建模基因表达语言分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。