arXiv:2409.09823cs.SDcs.MM2024-09被引 6

通过引入场景检测提升视频到音频生成的跨场景一致性。

Efficient Video to Audio Mapper with Visual Scene Detection

  • 在轻量级架构中加入视觉场景检测模块,增强多场景识别能力。
  • 在VGGSound数据集上,音画保真度与语义相关性均优于基线模型。
  • 适合需要高一致性音画生成的应用,如影视剪辑、智能创作。

视频到音频(V2A)生成旨在根据无声视频输入生成对应的音频。由于音视频特征具有跨模态性和序列性,该任务极具挑战性。近期工作在缩小音视频域差距方面取得显著进展,生成的音频能与视频内容语义对齐。然而,现有方法难以有效识别和处理视频中的多个场景,导致多场景切换时音频生成效果不佳。本文首先重构一个先进的V2A模型,采用轻量化改进架构,性能超越基线。随后提出一种新模型,引入场景检测模块以解决多场景切换问题。在VGGSound数据集上的实验表明,该模型能准确识别并处理视频中的多个场景,在音画保真度与语义相关性方面均优于基线模型。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have made significant progress in bridging the domain gap between video and audio, generating audio that is semantically aligned with the video content. However, a critical limitation of these approaches is their inability to effectively recognize and handle multiple scenes within a video, often leading to suboptimal audio generation in such cases. In this paper, we first reimplement a state-of-the-art V2A model with a slightly modified light-weight architecture, achieving results that outperform the baseline. We then propose an improved V2A model that incorporates a scene detector to address the challenge of switching between multiple visual scenes. Results on VGGSound show that our model can recognize and handle multiple scenes within a video and achieve superior performance against the baseline for both fidelity and relevance.

视频生成音画同步场景检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。