arXiv:2607.16789cs.LG2026-07

通过低秩分解解耦多模态信息,提升模型可解释性与泛化能力。

MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning

论文配图:MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning
图 1 · 摘自论文原文
  • 基于低秩适配思想,分离多模态共享与特异性信息
  • 在多个真实与模拟任务中实现准确的跨模态预测
  • 适合需要可控、可解释多模态表示的研究者使用

现实世界的感知与决策本质上是多模态的,需融合不同模态间的互补信号。然而,训练多模态模型面临两大挑战:一是大规模对齐多模态数据集难以获取,导致端到端训练困难;二是现有表示常将跨模态共享信息与模态特异性信息混杂,影响可解释性与控制能力。我们提出 MultiLoReFT,一种基于预训练单模态模型的高效、可扩展的低秩表示微调框架。该方法将低秩适配推广至多模态场景,学习可解释的投影子空间,实现共享与模态特异性信息的解耦。在模拟与真实世界基准测试中,其生成的表示既能支持多模态预测,又能明确揭示共享与特异性信息在各模态中的分布情况。

原文摘要 · Abstract (English)

Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned paired multimodal datasets is often impractical, making end-to-end multimodal training difficult. Second, existing multimodal representations frequently entangle information shared across modalities with modality-specific information, hindering interpretability and control. We introduce MultiLoReFT, an efficient and scalable low-rank representation fine-tuning framework for multimodal learning with pretrained unimodal models. MultiLoReFT extends low-rank adaptation to the multimodal setting and learns interpretable projection subspaces that decouple shared and modality-specific information. Across simulated and real-world benchmarks, it produces representations that support multimodal prediction while explicitly revealing how shared and modality-specific information is distributed across modalities.

多模态学习低秩微调可解释性信息解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。