发现并解决视频生成音频时凭空添加无视觉来源声音的问题
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
- 用多检测器投票框架系统评估音频幻觉现象
- 提出IH@vid和IH@dur指标,发现主流模型幻觉率超50%
- 提出HALCON方法,有效降低幻觉且不损害音画同步质量
视频到音频生成已取得显著进展,能自动为视频合成声音。但现有评估指标仅关注语义与时间对齐,忽视了关键缺陷:模型常生成无对应视觉源的声学事件,尤其语音和音乐。我们称此为插入幻觉(Insertion Hallucination),其根源是数据集偏差(如大量离屏声音),而当前指标完全无法检测。为此,我们构建了一个基于多音频事件检测器投票的系统评估框架,并引入两个新指标:IH@vid(含幻觉的视频比例)与IH@dur(幻觉持续时长占比)。在此基础上,提出HALCON方法:先生成初始音频暴露幻觉段,再识别并掩码不可靠视频特征,最后用修正后的条件重生成音频。在多个主流V2A基准测试中,结果显示顶尖模型存在严重幻觉;而HALCON平均将幻觉比例和时长降低超50%,且未损害甚至提升了传统音频质量与时间对齐指标。本工作首次正式定义、系统测量并有效缓解插入幻觉,推动更可靠、忠实的V2A模型发展。
原文摘要 · Abstract (English)
Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alignment, overlook a critical failure mode: models often generate acoustic events, particularly speech and music, that have no corresponding visual source. We term this phenomenon Insertion Hallucination and identify it as a systemic risk driven by dataset biases, such as the prevalence of off-screen sounds, that remains completely undetected by current metrics. To address this challenge, we first develop a systematic evaluation framework that employs a majority-voting ensemble of multiple audio event detectors. We also introduce two novel metrics to quantify the prevalence and severity of this issue: IH@vid (the fraction of videos with hallucinations) and IH@dur (the fraction of hallucinated duration). Building on this, we introduce HALCON to mitigate IH. HALCON follows a three-stage procedure: it first generates initial audio to expose hallucinated segments, then identifies and masks the corresponding unreliable video features, and finally regenerates the audio using the corrected conditioning. Experiments on several mainstream V2A benchmarks first reveal that state-of-the-art models suffer from severe IH. In contrast, our HALCON method reduces both the prevalence and duration of hallucinations by over 50\% on average, without degrading, and in some cases even improving, conventional metrics for audio quality and temporal synchronization. Our work is the first to formally define, systematically measure, and effectively mitigate Insertion Hallucination, paving the way for more reliable and faithful V2A models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。