arXiv:2505.22865cs.SDcs.AI2025-05ICML被引 10

用流匹配模型实现高质量、可实时播放的立体声语音合成

BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models

  • 将立体声渲染视为生成问题,采用条件流匹配建模
  • 实时推理下感知测试混淆率达42%,接近真实录音效果
  • 专为流式传输设计因果U-Net与连续推理管线

立体声渲染旨在基于单声道音频及说话人与听者位置,合成模仿自然听觉的立体声音频。尽管已有多种方法提出,但在渲染质量与流式推理方面仍存在挑战。生成与真实录音难以区分的高质量立体声音频,需精确建模立体声线索、房间混响与环境音。同时,实际应用要求流式推理。为此,我们提出基于流匹配的流式立体声语音合成框架BinauralFlow。将立体声渲染视为生成任务而非回归任务,设计条件流匹配模型以生成高质量音频;设计因果U-Net架构,仅依赖历史信息预测当前音频帧,适配生成模型的流式推理;引入包含流式STFT/ISTFT、缓冲池、中点求解器与早期跳过调度的连续推理流水线,提升渲染连贯性与速度。定量与定性评估表明,本方法优于现有最优技术。感知实验进一步显示,模型输出几乎无法与真实录音区分,混淆率达42%。

原文摘要 · Abstract (English)

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from real-world recordings, with a $42\%$ confusion rate.

语音合成立体声渲染流式推理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。