从无声视频生成随视角变化的空间音频,提升沉浸感。
ViSAGe: Video-to-Spatial Audio Generation
- 用视觉特征引导音频生成,端到端合成空间音效。
- 在102K视频数据集上实现高质量音效,优于两阶段方法。
- 适合虚拟现实、影视制作等需要动态空间音效的场景。
空间音频对提升音视频沉浸感至关重要,但传统制作需复杂设备与专业技能。本文提出从无声视频直接生成一阶全向声(first-order ambisonics)的新任务,构建了包含102,000段5秒YouTube视频及其对应一阶全向声的YT-Ambigen数据集,并设计基于音频能量图与显著性度量的新评估指标。提出端到端框架ViSAGe,利用CLIP视觉特征和自回归音频编码器,结合方向与视觉引导生成空间音频。实验表明,ViSAGe生成的音频在时序上对齐且质量高,能随视角变化适应,优于先生成音频再做空间化的两阶段方法。定性结果验证其可生成与视频内容一致的动态空间音效。
原文摘要 · Abstract (English)
Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem of generating first-order ambisonics, a widely used spatial audio format, directly from silent videos. To support this task, we introduce YT-Ambigen, a dataset comprising 102K 5-second YouTube video clips paired with corresponding first-order ambisonics. We also propose new evaluation metrics to assess the spatial aspect of generated audio based on audio energy maps and saliency metrics. Furthermore, we present Video-to-Spatial Audio Generation (ViSAGe), an end-to-end framework that generates first-order ambisonics from silent video frames by leveraging CLIP visual features, autoregressive neural audio codec modeling with both directional and visual guidance. Experimental results demonstrate that ViSAGe produces plausible and coherent first-order ambisonics, outperforming two-stage approaches consisting of video-to-audio generation and audio spatialization. Qualitative examples further illustrate that ViSAGe generates temporally aligned high-quality spatial audio that adapts to viewpoint changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。