自动分析多模态假信息数据集中每条样本的模态偏见,揭示检测器依赖单一模态的问题。
Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- 提出三种粒度不同的模态偏见量化方法,从粗到细分析信息依赖关系。
- 实验发现融合多视角分析能提升可靠性,但检测器会引入波动偏差。
- 在平衡样本上不同方法一致,在偏见样本上分歧明显,适合研究假信息检测漏洞。
多个多模态假信息基准数据集存在特定模态偏向,使检测器仅依赖单一模态即可预测。以往研究多在数据集层面量化偏见或手动识别模态与标签间的虚假关联,缺乏样本级洞察且难以扩展至海量网络信息。本文提出面向样本级的自动化模态偏见识别设计,基于不同粒度理论提出三种量化方法:1)粗粒度模态收益评估;2)中粒度信息流量化;3)细粒度因果分析。通过在两个主流基准上的人工评估验证有效性。实验揭示三个重要发现:1)融合多视角对可靠自动化分析至关重要;2)自动化分析易受检测器影响产生波动;3)不同视角在模态平衡样本上一致性高,但在偏见样本上分歧显著,为未来研究提供方向。
原文摘要 · Abstract (English)
Numerous multimodal misinformation benchmarks exhibit bias toward specific modalities, allowing detectors to make predictions based solely on one modality. While previous research has quantified bias at the dataset level or manually identified spurious correlations between modalities and labels, these approaches lack meaningful insights at the sample level and struggle to scale to the vast amount of online information. In this paper, we investigate the design for automated recognition of modality bias at the sample level. Specifically, we propose three bias quantification methods based on theories/views of different levels of granularity: 1) a coarse-grained evaluation of modality benefit; 2) a medium-grained quantification of information flow; and 3) a fine-grained causality analysis. To verify the effectiveness, we conduct a human evaluation on two popular benchmarks. Experimental results reveal three interesting findings that provide potential direction toward future research: 1)~Ensembling multiple views is crucial for reliable automated analysis; 2)~Automated analysis is prone to detector-induced fluctuations; and 3)~Different views produce a higher agreement on modality-balanced samples but diverge on biased ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。