直接从视频生成沉浸式立体音频,解决传统方法音画不同步问题。
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
- 端到端框架,双分支建模音频潜变量流,避免分步处理误差累积。
- 在97K视频-立体音频对上训练,实现视角变化下精准声场适配。
- 适合虚拟现实、影视制作等需真实空间听感的场景应用。
尽管视频转音频技术取得进展,但多数工作仍聚焦于单声道输出,缺乏空间沉浸感。现有立体音频方法受限于先生成单声道再进行空间化的两阶段流程,常导致误差累积和时空不一致。为此,本文提出直接从无声视频生成端到端立体空间音频的新任务。为支持该任务,我们构建了包含约97,000组视频-立体音频对的BiAudio数据集,覆盖多样真实场景与摄像机旋转轨迹,通过半自动化管道生成。同时提出ViSAudio框架,采用条件流匹配与双分支音频生成结构,两个专用分支分别建模音频潜变量流。结合条件时空模块,平衡通道间一致性并保留独特空间特征,确保音频与输入视频的精确时空对齐。大量实验表明,ViSAudio在客观指标与主观评测中均优于现有最先进方法,能生成高质量、具备空间沉浸感的立体音频,有效适应视角变化、声源运动及多样化声学环境。
原文摘要 · Abstract (English)
Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-stage pipeline that first generates mono audio and then performs spatialization, often resulting in error accumulation and spatio-temporal inconsistencies. To address this limitation, we introduce the task of end-to-end binaural spatial audio generation directly from silent video. To support this task, we present the BiAudio dataset, comprising approximately 97K video-binaural audio pairs spanning diverse real-world scenes and camera rotation trajectories, constructed through a semi-automated pipeline. Furthermore, we propose ViSAudio, an end-to-end framework that employs conditional flow matching with a dual-branch audio generation architecture, where two dedicated branches model the audio latent flows. Integrated with a conditional spacetime module, it balances consistency between channels while preserving distinctive spatial characteristics, ensuring precise spatio-temporal alignment between audio and the input video. Comprehensive experiments demonstrate that ViSAudio outperforms existing state-of-the-art methods across both objective metrics and subjective evaluations, generating high-quality binaural audio with spatial immersion that adapts effectively to viewpoint changes, sound-source motion, and diverse acoustic environments. Project website: https://kszpxxzmc.github.io/ViSAudio-project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。