无配对多模态数据中分离共享成分,突破传统配对假设限制。
Identifiable Shared Component Analysis of Unpaired Multimodal Mixtures
- 基于分布差异最小化设计新损失函数,实现无配对数据下的共享成分识别
- 在合成与真实数据上验证了共享成分可识别性,条件比传统方法更宽松
- 适用于语音-文本等无配对多模态场景,适合做跨模态表示学习的研究者
多模态学习的核心任务是整合不同特征空间(如文本与音频)的信息,获得对模态不变的数据本质表示。已有研究表明,在跨模态样本配对的条件下,经典工具如典型相关分析(CCA)可证明地识别出共享成分,仅存在微小歧义。本文进一步研究无配对多模态线性混合数据中的共享成分可识别性问题。提出一种基于分布差异最小化的损失函数,并推导出一系列保证共享成分可识别的充分条件。这些条件基于跨模态分布差异刻画和保密度变换消除,比依赖独立成分分析的方法更为宽松。通过引入合理结构约束,还提供了更宽松的识别条件,该设计受实际应用中可用辅助信息的启发。识别结论在合成数据与真实世界数据上均得到充分验证。
原文摘要 · Abstract (English)
A core task in multi-modal learning is to integrate information from multiple feature spaces (e.g., text and audio), offering modality-invariant essential representations of data. Recent research showed that, classical tools such as {\it canonical correlation analysis} (CCA) provably identify the shared components up to minor ambiguities, when samples in each modality are generated from a linear mixture of shared and private components. Such identifiability results were obtained under the condition that the cross-modality samples are aligned/paired according to their shared information. This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned. A distribution divergence minimization-based loss is proposed, under which a suite of sufficient conditions ensuring identifiability of the shared components are derived. Our conditions are based on cross-modality distribution discrepancy characterization and density-preserving transform removal, which are much milder than existing studies relying on independent component analysis. More relaxed conditions are also provided via adding reasonable structural constraints, motivated by available side information in various applications. The identifiability claims are thoroughly validated using synthetic and real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。