构建新基准,系统性破坏多模态一致性,检验模型可靠性与判断犹豫能力。
Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention
- 通过逐步破坏视频、音频、文本三模态信息,分离各模态贡献。
- 模型在三模态全毁时仍保持60%-100%高信心,严重低估风险。
- 适合评估多模态系统对矛盾信息的应对能力与可信判断机制。
现有跨模态基准难以区分模态依赖与信息不对称,因自然共现的模态具有相关但不等的信息。本文提出OMD-Bench,初始时三模态(视频、音频、文本)均呈现同一可独立感知的锚点,随后系统性地施加八种损坏条件,共4,080个样本,覆盖27个锚点。评估十种跨模态模型在零样本与思维链提示下的表现,发现当两模态受损时模型过度弃权,而三模态全损时却严重低估风险,自信度仍高达60%-100%。思维链提示虽提升弃权与人类判断的一致性,却加剧了过度自信。OMD-Bench为诊断跨模态系统中的模态依赖、鲁棒性及不确定性校准提供了可解释的诊断工具。
原文摘要 · Abstract (English)
Existing omni-modal benchmarks attempt to measure modality-specific contributions, but their measurements are confounded: naturally co-occurring modalities carry correlated yet unequal information, making it unclear whether results reflect true modality reliance or information asymmetry. We introduce OMD-Bench, where all modalities are initially congruent - each presenting the same anchor, an object or event independently perceivable through video, audio, and text - which we then systematically corrupt to isolate each modality's contribution. We also evaluate calibrated abstention: whether models appropriately refrain from answering when evidence is conflicting. The benchmark comprises 4,080 instances spanning 27 anchors across eight corruption conditions. Evaluating ten omni-modal models under zero-shot and chain-of-thought prompting, we find that models over-abstain when two modalities are corrupted yet under-abstain severely when all three are, while maintaining high confidence (~60-100%) even under full corruption. Chain-of-thought prompting improves abstention alignment with human judgment but amplifies overconfidence rather than mitigating it. OMD-Bench provides a diagnostic benchmark for diagnosing modality reliance, robustness to cross-modal inconsistency, and uncertainty calibration in omni-modal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。