评测多模态模型在跨模态推理中的融合与抗冲突能力,揭示其易受单一模态主导的缺陷。
C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

- 构建包含4类模态的3404样本基准,评估信息融合与矛盾化解能力
- 顶尖模型准确率仅73.17%,远低于人类88.64%,开源模型在冲突下崩溃
- 发现86-95%失败源于模态霸权,文本注意力占比超87%,需持续跨模态探索
当前多模态大模型虽可处理多样感官输入,但推理仍严重偏向主导模态,导致跨模态推理脆弱。我们提出C³PO,一个包含3,404个样本的基准,涵盖视频、音频、图像和文本,用于评估两种能力:信息组合(融合分散证据)与反事实冲突(化解故意矛盾)。其成对的IC/CC结构与四层设计可精准诊断跨模态推理失败的时间与原因。该基准通过25个逻辑基础模板的全自动流程构建。结果显示,人类准确率达88.64%,最佳模型(Gemini-3.1-Pro)仅达73.17%,开源模型在冲突下性能崩溃。注意力探测显示,86%-95%的失败源于模态主导:模型固守某一模态而忽略矛盾证据,87%-95%注意力集中于文本。中层注意力熵可预测正确性:持续探索者成功,过早坍缩者失败。同等复杂度模板间56分准确率差距表明,表现取决于模态在冲突解决中的结构角色,而非模态组合。研究揭示,多模态感知不等于鲁棒推理;架构必须支持持续的跨模态注意力以避免过早定论。
原文摘要 · Abstract (English)
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。