arXiv:2509.20971cs.SDcs.AI2025-09中稿 · AIML Systems 2025被引 3

优化语音对话系统延迟,提升实时交互体验。

i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents

  • 采用端到端架构,结合音频与文本上下文生成自然语音。
  • 减少TTS解码中RVQ迭代次数可显著降低延迟。
  • 适合需要低延迟语音交互的智能助手场景。

我们实验了一种低延迟、端到端的语音到语音(V-2-V)通信模型,旨在优化实时对话应用。通过分析自动语音识别(ASR)、文本转语音(TTS)和对话管理等关键组件,研究如何在保持高质量交互的前提下减少处理时间,从而识别V-2-V系统的优化杠杆。结果表明,生成富有情感、包含自然停顿和感叹的逼真语音的TTS组件对实时因子(RTF)影响最大。所实验的V-2-V架构采用CSM1b,能够理解对话中的语调与上下文,通过融合先前交流的音频与文本生成语义准确的语音。我们探索了在TTS解码器中减少残差向量量化(RVQ)迭代次数的优化方案,尽管会轻微降低语音质量。实验还表明,基于CSM的V-2-V系统中,最重要且有效的优化手段是减少RVQ迭代次数及Mimi中使用的码本数量。

原文摘要 · Abstract (English)

We experiment with a low-latency, end-to-end voice-to-voice communication model to optimize it for real-time conversational applications. By analyzing components essential to voice to voice (V-2-V) system viz. automatic speech recognition (ASR), text-to-speech (TTS), and dialog management, our work analyzes how to reduce processing time while maintaining high-quality interactions to identify the levers for optimizing V-2-V system. Our work identifies that TTS component which generates life-like voice, full of emotions including natural pauses and exclamations has highest impact on Real time factor (RTF). The experimented V-2-V architecture utilizes CSM1b has the capability to understand tone as well as context of conversation by ingesting both audio and text of prior exchanges to generate contextually accurate speech. We explored optimization of Residual Vector Quantization (RVQ) iterations by the TTS decoder which come at a cost of decrease in the quality of voice generated. Our experimental evaluations also demonstrate that for V-2-V implementations based on CSM most important optimizations can be brought by reducing the number of RVQ Iterations along with the codebooks used in Mimi.

语音生成低延迟对话系统TTS优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。