通过结构扰动测试多模态模型在信息失序时的可靠性。
Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

- 固定问题与答案空间,对文本、视觉、音频分别施加结构扰动
- 结构损坏使准确率下降,图文损坏形成最稳定的脆弱边界
- 多模态退化非简单叠加,适合评估模型真实跨模态能力
多模态大模型常在干净的文本-视觉-音频输入上评估,各通道完整、同步且可读。这类得分常被视为跨模态融合鲁棒性的证据,但干净评估无法判断成功是依赖稳定跨模态结构,还是仅依赖完整输入中的线索。为此,我们定义了“模态故障线”:当某模态仍存在且人类可解读,但其内部证据结构被扰动时,模型行为变得不稳定的边界。我们提出SCEval(结构扰动评估)诊断协议,在保持问题、答案空间和模态通道不变的前提下,对文本、视觉、音频单独或联合施加受控结构扰动。数据基于来自Social-IQ、OmniBench和VALOR的273个经人工验证的三模态样本,评估15个专有及开源多模态系统。结果表明,结构扰动降低原始准确率,文本-视觉损坏形成最稳定的共享故障线,多模态退化具有非叠加性而非单纯由受损模态数量决定。因此,干净环境下的多模态准确率无法证明模型在跨模态证据结构不可靠时仍能保持可靠。
原文摘要 · Abstract (English)
Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tell whether success depends on stable cross-modal structure or on cues sufficient only in intact inputs. To address this gap, we define a modality fault line: a boundary at which model behavior becomes unstable when a modality remains present and human-interpretable, but its internal evidence structure is perturbed. We introduce SCEval (Structure-Corruption Evaluation) a diagnostic evaluation protocol that keeps the question, answer space, and modality channels fixed while applying controlled structural corruptions to text, vision, and audio individually and jointly. Built from $273$ human-verified tri-modal examples from Social-IQ, OmniBench, and VALOR, SCEval evaluates $15$ proprietary and open-source omni-modal systems. The results show that structural corruption lowers clean accuracy, text--vision damage forms the most stable shared fault line, and multi-modal degradation is non-additive rather than a simple function of the number of corrupted modalities. Clean omni-modal accuracy therefore does not establish that a model will remain reliable when cross-modal evidence becomes structurally unreliable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。