arXiv:2503.01728cs.LGstat.ME2025-03被引 1

提出深度充分模态学习框架,自动选最优模态组合以提升效率

DeepSuM: Deep Sufficient Modality Learning Framework

  • 独立学习各模态表征,按自身空间评估重要性
  • 实现模态融合与选择的联合优化,提升系统效率
  • 适合资源受限场景下的多模态应用开发

多模态学习已成为构建鲁棒学习模型的关键方法,广泛应用于多媒体、机器人、大语言模型和医疗等领域。由于不同模态的成本和资源消耗差异显著,多模态系统的效率成为关键挑战。这要求有效的模态选择机制,在性能提升与资源开销间取得平衡。本文提出一种新的模态选择框架,独立学习各模态的表示,使其在各自表征空间中评估重要性,支持定制化编码器设计,并促进具有不同特性的模态联合分析。该框架通过优化模态整合与选择,旨在提升多模态学习的效率与效果。

原文摘要 · Abstract (English)

Multimodal learning has become a pivotal approach in developing robust learning models with applications spanning multimedia, robotics, large language models, and healthcare. The efficiency of multimodal systems is a critical concern, given the varying costs and resource demands of different modalities. This underscores the necessity for effective modality selection to balance performance gains against resource expenditures. In this study, we propose a novel framework for modality selection that independently learns the representation of each modality. This approach allows for the assessment of each modality's significance within its unique representation space, enabling the development of tailored encoders and facilitating the joint analysis of modalities with distinct characteristics. Our framework aims to enhance the efficiency and effectiveness of multimodal learning by optimizing modality integration and selection.

多模态学习模态选择表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。