用混响语音生成真实感混响脉冲响应,支持可控声学模拟。
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
- 基于扩散变换器,从混响语音预测复频谱混响脉冲响应。
- 生成的脉冲响应更丰富,且在标准指标上表现优异。
- 可与现有语音流程模块化集成,适合可控声学仿真任务。
盲区脉冲响应(RIR)估计是捕捉和传递声学特性的重要任务;然而现有方法常受限于建模能力,并在未见条件下性能下降。此外,新兴的生成式音频应用需要更灵活的脉冲响应生成方法。我们提出 Gencho,一种基于扩散变换器的模型,能够从混响语音中预测复频谱 RIR。结构感知编码器利用早期与晚期反射之间的分离性,将输入音频编码为鲁棒表征以实现条件控制,而扩散解码器则从中生成多样且听觉真实的脉冲响应。Gencho 可模块化集成至标准语音处理流程中,实现声学匹配。实验表明,其生成的 RIR 比非生成基线更丰富,同时在标准 RIR 指标上保持强劲表现。我们进一步展示了其在文本条件下的 RIR 生成应用,凸显 Gencho 在可控声学仿真与生成式音频任务中的多功能性。
原文摘要 · Abstract (English)
Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover, emerging generative audio applications call for more flexible impulse response generation methods. We propose Gencho, a diffusion-transformer-based model that predicts complex spectrogram RIRs from reverberant speech. A structure-aware encoder leverages isolation between early and late reflections to encode the input audio into a robust representation for conditioning, while the diffusion decoder generates diverse and perceptually realistic impulse responses from it. Gencho integrates modularly with standard speech processing pipelines for acoustic matching. Results show richer generated RIRs than non-generative baselines while maintaining strong performance in standard RIR metrics. We further demonstrate its application to text-conditioned RIR generation, highlighting Gencho's versatility for controllable acoustic simulation and generative audio tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。