让视觉别抢戏,用声音对比优化模型听觉判断
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
- 引入输出与输入双对比机制,惩罚视觉误导的音频生成
- 在AVQA和VQA任务上显著降低音频幻觉,提升听觉对齐精度
- 适合关注多模态可靠性、避免视觉偏见的研究者
尽管音频-视觉语言模型(AVLMs)近年取得显著进展,其可靠性仍受限于跨模态幻觉。一种普遍现象是视频驱动的音频幻觉:模型常依赖视觉线索生成预期声音,忽略真实听觉信息。为破解这种深层的视觉主导问题,本文提出音频对比偏好优化(ACPO)。该双轴偏好学习框架包含输出对比目标,惩罚将视觉描述伪装成音频事实的行为;以及输入对比目标,通过交换音频轨道,显式惩罚对真实听觉信号不敏感的生成。大量实验表明,ACPO显著提升了音频的忠实性,有效缓解了音频幻觉。
原文摘要 · Abstract (English)
While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal hallucination. A particularly pervasive manifestation is video-driven audio hallucination: models routinely exploit visual shortcuts to hallucinate expected sounds, discarding true auditory evidence. To counteract this deeply ingrained visual dominance, we propose Audio-Contrastive Preference Optimization (ACPO). This dual-axis preference learning framework introduces an output-contrastive objective to penalize visual descriptions masquerading as audio facts, alongside an input-contrastive objective that swaps audio tracks to explicitly penalize generation invariant to the true auditory signal. Extensive experiments demonstrate that ACPO establishes highly faithful audio grounding and mitigates audio hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。