系统梳理多模态对齐与融合方法,助力跨模态智能发展
Multimodal Alignment and Fusion: A Survey
- 按数据、特征、输出层级划分对齐融合结构,归纳七类方法
- 综述260+研究,揭示跨模态错位与计算瓶颈等核心挑战
- 适合多模态学习、具身智能等领域的研究人员参考
本综述全面回顾了机器学习领域中多模态对齐与融合的最新进展,得益于文本、图像、音频和视频等多源数据日益丰富。不同于以往聚焦单一模态或有限融合策略的综述,本文提出以结构为中心、以方法为导向的框架,强调可泛化技术。通过数据级、特征级、输出级融合的结构视角,以及统计、核方法、图模型、生成模型、对比学习、注意力机制和大语言模型(LLM)等方法范式,系统分类并分析关键方法,基于对260余篇相关研究的广泛调研。同时,本文指出跨模态错位、计算瓶颈、数据质量及模态鸿沟等关键挑战,并介绍近期应对措施。应用涵盖社交媒体分析、医学影像、情感识别和具身智能等领域,展示鲁棒多模态系统的现实影响。研究成果旨在为未来研究提供指导,推动多模态学习系统在可扩展性、鲁棒性和跨领域泛化方面的优化。
原文摘要 · Abstract (English)
This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio, and video. Unlike previous surveys that often focus on specific modalities or limited fusion strategies, our work presents a structure-centric and method-driven framework that emphasizes generalizable techniques. We systematically categorize and analyze key approaches to alignment and fusion through both structural perspectives -- data-level, feature-level, and output-level fusion -- and methodological paradigms -- including statistical, kernel-based, graphical, generative, contrastive, attention-based, and large language model (LLM)-based methods, drawing insights from an extensive review of over 260 relevant studies. Furthermore, this survey highlights critical challenges such as cross-modal misalignment, computational bottlenecks, data quality issues, and the modality gap, along with recent efforts to address them. Applications ranging from social media analysis and medical imaging to emotion recognition and embodied AI are explored to illustrate the real-world impact of robust multimodal systems. The insights provided aim to guide future research toward optimizing multimodal learning systems for improved scalability, robustness, and generalizability across diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。