arXiv:2511.22826cs.CV2025-11被引 5

研究多模态模型在信息冲突时的表现,提出增强其跨模态推理可靠性的方法。

Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs

  • 构建MMA-Bench评测基准,测试模型对音视频不一致等冲突信息的反应。
  • 发现当前多模态大模型在音视频错配和误导文本下表现脆弱。
  • 提出模态对齐调优策略,让模型学会合理判断各模态的可信度。

尽管多模态大语言模型(MLLMs)取得显著进展,但一个根本问题仍未解决:它们是否对矛盾模态具有鲁棒性?为严谨研究此问题,我们提出了MMA-Bench,包含视频与任务,用于探测模型对特定模态的依赖程度。通过黑盒与白盒可解释性技术,我们对开源与闭源的MLLMs进行了批判性分析。结果表明,当前模型在音视频错配及简单误导文本下表现不佳,缺乏稳健的多模态推理能力。基于这些发现,我们提出一种模态对齐调优策略,教会模型在何时优先、利用或忽略特定模态线索。通过大量实验与分析,证明该调优方法显著提升了多模态定位能力。本工作提供了可解释性工具与构建内在可靠跨模态推理模型的明确路径。代码与数据集将公开。

原文摘要 · Abstract (English)

Despite remarkable advancements in Multimodal Large Language Models (MLLMs), a fundamental question remains: are MLLMs robust to contradicting modalities? To rigorously study this, we introduce MMA-Bench comprising videos and tasks that probe a model's reliance on specific modalities. Using black-box and white-box interpretability techniques, we provide a critical analysis of the brittleness of both open- and closed-sourced MLLMs. We show that current MLLMs struggle under misaligned audio-visual pairs and simple misleading text, thereby lacking robust multi-modal reasoning. Building on these findings, we propose a modality alignment tuning strategy to teach the model when to prioritize, leverage, or ignore specific modality cues. Through extensive experiments and analysis, we show that our alignment tuning yields demonstrably stronger multimodal grounding. This work provides both interpretability tools and a clear path toward developing MLLMs with intrinsically reliable cross-modal reasoning. Code and dataset will be publicly available.

多模态模型鲁棒性可解释性对齐调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。