arXiv:2606.05109cs.LG2026-06

突破双模态限制,实现多模态表征解耦的高效框架

RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities

论文配图:RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities
图 1 · 摘自论文原文
  • 基于插件式架构直接处理预提取嵌入,无需联合预训练
  • 统一优化共享与独有成分,在多模态任务中保持性能并降低计算开销
  • 理论保证解耦最优性,适用于任意模态和基础模型

为充分发挥多模态数据潜力,需超越现有对齐与融合方法,充分挖掘所有跨模态交互关系,同时保留模态特异性信息。学习解耦表征是识别观测数据中隐藏的共享与独特因素的合理途径。然而,尽管多模态解耦具有吸引力,现有方法因固有的可扩展性瓶颈,大多局限于双模态场景。为此,我们提出RePercENT,一种自监督框架,旨在突破此类限制,实现超过两模态的可扩展成对解耦。通过多模态‘即插即用’架构,该方法直接作用于预提取嵌入,无需复杂联合预训练,且不对底层模态或基础模型主干做假设。此外,我们引入联合优化目标,同时推导共享与独有成分,并提供形式化理论保证,刻画解耦方案的最优性。在多种模态和任务上,RePercENT成功恢复解耦成分,保持竞争力性能,显著降低计算复杂度。

原文摘要 · Abstract (English)

To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data. However, while multimodal disentanglement is a compelling paradigm, existing methods are largely confined to the two-modality regime due to its inherent scalability bottleneck. To address this, we propose RePercENT, a self-supervised framework designed to surpass these limitations and unlocks scalable pairwise disentanglement beyond two modalities. Through a multimodal `plug-and-play' architecture, our approach operates directly on pre-extracted embeddings, eliminating the need for extensive joint pre-training while making no assumptions regarding the underlying modalities or foundation model backbones. Moreover, we introduce a joint optimization objective for simultaneously deriving the shared and unique components, and provide formal theoretical guarantees that characterize the optimality of our solution. Across diverse modalities and tasks, RePercENT successfully recovers disentangled components while maintaining competitive performance and significantly reducing computational complexity.

多模态解耦表征自监督可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。