融合多模态结构数据,用新模型发现共享因子并提升可解释性。
Factor Analysis with Correlated Topic Model for Multi-Modal Data
- 结合因子分析与相关主题模型,构建多视角多结构贝叶斯框架。
- 在文本、视频、音乐和新冠数据上,识别结构化数据聚类更准确。
- 通过因子旋转增强二值特征的可解释性,适合生物、媒体分析场景。
融合多种数据模态能揭示底层现象的深层见解。多模态因子分析(FA)可发现不同简单模态共享的变异轴,每个样本以特征向量表示。然而,传统FA不适用于具有聚类结构的结构化数据,如文本或单细胞测序数据,其中每样本包含多个数据点。为此,我们提出FACTM——一种新颖的多视图、多结构贝叶斯模型,结合了因子分析与相关主题建模,并采用变分推断进行优化。此外,我们提出一种因子旋转方法,以提升对二值特征的可解释性。在文本与视频基准数据集,以及真实世界中的音乐和新冠数据集上,FACTM在识别结构化数据聚类方面优于现有方法,并能通过推断共享且可解释的因子,有效整合结构化与简单模态数据。
原文摘要 · Abstract (English)
Integrating various data modalities brings valuable insights into underlying phenomena. Multimodal factor analysis (FA) uncovers shared axes of variation underlying different simple data modalities, where each sample is represented by a vector of features. However, FA is not suited for structured data modalities, such as text or single cell sequencing data, where multiple data points are measured per each sample and exhibit a clustering structure. To overcome this challenge, we introduce FACTM, a novel, multi-view and multi-structure Bayesian model that combines FA with correlated topic modeling and is optimized using variational inference. Additionally, we introduce a method for rotating latent factors to enhance interpretability with respect to binary features. On text and video benchmarks as well as real-world music and COVID-19 datasets, we demonstrate that FACTM outperforms other methods in identifying clusters in structured data, and integrating them with simple modalities via the inference of shared, interpretable factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。