实时生成高保真空间音频,支持视频与文本输入。
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

- 采用因果自回归扩散变压器架构,实现流式音频生成。
- 在多模态数据上达到高空间感知精度,延迟低。
- 适合沉浸式音视频应用,如VR/AR与元宇宙场景。
实时精准的空间音频生成对营造沉浸式体验至关重要。然而,现有技术常在生成质量与高推理延迟之间存在权衡,且难以从多模态输入中准确捕捉空间信息。为此,我们提出SwanSphere——一种统一的流式框架,可基于全景视频和文本提示生成高保真空间音频。主要贡献包括:1)提出因果自回归扩散变压器架构,支持高质量流式音频生成;2)设计空间视频-音频对比(SVAC)学习策略,对齐视频编码器与声学域,并引入多目标在线直接偏好优化(ODPO)方案,显著提升空间感知能力与多模态合成鲁棒性;3)为缓解空间音频数据集稀缺问题,构建自动化标注流水线以生成详细空间描述。实验表明,SwanSphere在视频到空间音频及文本到空间音频任务中均表现优异。演示可访问:https://swanaigc.github.io。
原文摘要 · Abstract (English)
Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial information from multimodal inputs. To address these challenges, we propose SwanSphere, a unified streaming framework for high-fidelity spatial audio generation from panoramic videos and text prompts. SwanSphere mainly makes the following contributions: 1) We introduce a causal autoregressive diffusion transformer architecture that enables streaming high-quality spatial audio generation. 2) We design a Spatial Video-Audio Contrastive (SVAC) learning strategy to align the video encoder with the acoustic domain, and further employ a multi-objective online direct preference optimization (ODPO) scheme, resulting in strong spatial perception and robust multimodal spatial audio synthesis. 3) To alleviate the current scarcity of spatial audio datasets, we also develop an automated annotation pipeline for generating detailed spatial captions. Experimental results demonstrate that SwanSphere achieves superior performance in both video-to-spatial and text-to-spatial audio generation tasks. Demos can be found at: https://swanaigc.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。