发现视觉语音描述模型严重依赖音频,提出新数据集缓解偏见
Listening without Looking: Modality Bias in Audio-Visual Captioning
- 通过破坏音视频流测试模型鲁棒性,发现模型严重偏听
- 在新数据集上训练后,模型对双模态的依赖更均衡
- 适合研究多模态融合偏见与公平性的学者参考
音视频描述任务旨在联合建模声音与视觉信息生成整体场景描述。尽管近期方法通过复杂的模态融合提升了性能,但当前模型中两种模态的互补程度以及在单模态退化时的鲁棒性仍不明确。本文针对最先进的音视频描述模型LAVCap,系统性地进行模态鲁棒性测试,通过选择性抑制或破坏音频或视觉流,量化其敏感性和互补性。分析显示,LAVCap对音频流存在显著偏倚。为评估模型对双模态的平衡使用,本文在AudioCaps基础上新增同时描述音视频内容的文本标注,构建AudioVisualCaps数据集。实验报告了LAVCap在AudioVisualCaps上的基线结果,并在该数据集上进行模态鲁棒性测试,结果显示:在AudioVisualCaps上训练的LAVCap比在AudioCaps上训练的模型表现出更少的模态偏倚。
原文摘要 · Abstract (English)
Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two modalities are complementary in current audio-visual captioning models and how robust these models are when one modality is degraded. We address these questions by conducting systematic modality robustness tests on LAVCap, a state-of-the-art audio-visual captioning model, in which we selectively suppress or corrupt the audio or visual streams to quantify sensitivity and complementarity. The analysis reveals a pronounced bias toward the audio stream in LAVCap. To evaluate how balanced audio-visual captioning models are in their use of both modalities, we augment AudioCaps with textual annotations that jointly describe the audio and visual streams, yielding the AudioVisualCaps dataset. In our experiments, we report LAVCap baseline results on AudioVisualCaps. We also evaluate the model under modality robustness tests on AudioVisualCaps and the results indicate that LAVCap trained on AudioVisualCaps exhibits less modality bias than when trained on AudioCaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。