arXiv:2507.20089cs.LGstat.ME2025-07被引 4

提出统一多模态融合框架,通过模型互学提升预测性能。

Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning

  • 构建多模型协同群体,基于不同模态组合的隐表示进行融合。
  • 在模拟与真实数据上均显著优于传统早/中/晚融合方法。
  • 适用于医疗诊断等需多模态融合的场景,兼容各类模型。

开发高效的多模态数据融合策略对提升统计机器学习在自动驾驶、医学诊断等领域的预测能力至关重要。传统融合方法(早、中、晚融合)各有优劣。本文提出Meta Fusion,一个统一现有策略的灵活框架。受深度互学习和集成学习启发,该框架基于多模态隐表示的不同组合构建模型群体,并通过群体内软信息共享进一步提升性能。该方法在学习隐表示时与模型无关,可灵活适应各模态特性。理论上,软信息共享机制可降低泛化误差;实验上,Meta Fusion在广泛仿真研究中持续优于传统融合策略,并在阿尔茨海默病检测和神经解码等真实应用中得到验证。

原文摘要 · Abstract (English)

Developing effective multimodal data fusion strategies has become increasingly essential for improving the predictive power of statistical machine learning methods across a wide range of applications, from autonomous driving to medical diagnosis. Traditional fusion methods, including early, intermediate, and late fusion, integrate data at different stages, each offering distinct advantages and limitations. In this paper, we introduce Meta Fusion, a flexible and principled framework that unifies these existing strategies as special cases. Motivated by deep mutual learning and ensemble learning, Meta Fusion constructs a cohort of models based on various combinations of latent representations across modalities, and further boosts predictive performance through soft information sharing within the cohort. Our approach is model-agnostic in learning the latent representations, allowing it to flexibly adapt to the unique characteristics of each modality. Theoretically, our soft information sharing mechanism reduces the generalization error. Empirically, Meta Fusion consistently outperforms conventional fusion strategies in extensive simulation studies. We further validate our approach on real-world applications, including Alzheimer's disease detection and neural decoding.

多模态融合互学习模型泛化医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。