发现多模态模型因单模态错误误导整体判断,提出诊断方法
When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning
- 将各模态视为独立代理,自动生成预测与评估
- 识别出主导错误的'破坏者'模态,揭示融合缺陷
- 适用于情感识别等任务,助于定位模型失败原因
尽管多模态大语言模型发展迅速,其推理过程仍不透明:难以确定哪一模态主导预测、冲突如何解决,或何时某一模态占据主导。本文提出‘模态破坏’这一诊断性故障模式,即高置信度的单模态错误会覆盖其他证据,误导融合结果。为此,我们设计了一种轻量级、模型无关的评估层,将每个模态视为独立代理,生成候选标签和简短自我评估,用于审计。通过简单融合机制聚合输出,可暴露支持正确结果的贡献者与误导结果的破坏者。在基础模型的多模态情感识别基准上进行案例研究,揭示了系统性的可靠性特征,有助于区分失败是源于数据集偏差还是模型局限。该框架为多模态推理提供了诊断基础,支持对融合机制的严谨审计,并指导潜在干预措施。
原文摘要 · Abstract (English)
Despite rapid growth in multimodal large language models (MLLMs), their reasoning traces remain opaque: it is often unclear which modality drives a prediction, how conflicts are resolved, or when one stream dominates. In this paper, we introduce modality sabotage, a diagnostic failure mode in which a high-confidence unimodal error overrides other evidence and misleads the fused result. To analyze such dynamics, we propose a lightweight, model-agnostic evaluation layer that treats each modality as an agent, producing candidate labels and a brief self-assessment used for auditing. A simple fusion mechanism aggregates these outputs, exposing contributors (modalities supporting correct outcomes) and saboteurs (modalities that mislead). Applying our diagnostic layer in a case study on multimodal emotion recognition benchmarks with foundation models revealed systematic reliability profiles, providing insight into whether failures may arise from dataset artifacts or model limitations. More broadly, our framework offers a diagnostic scaffold for multimodal reasoning, supporting principled auditing of fusion dynamics and informing possible interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。