arXiv:2510.16156eess.AScs.AI2025-10中稿 · the IEEE ASRU 2025…被引 2

让大模型推理过程实时语音化,支持用户随时打断提问。

AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning

  • 异步架构分离语音前端与推理后端,实现思考与播报并行。
  • 交互延迟降低600倍以上,任务准确率保持竞争力。
  • 适合需要实时干预的高风险决策场景,如医疗或金融。

复杂推理任务中的人机协作需要用户理解并参与模型的思考过程,而传统链式思维(CoT)的单一文本输出难以实现。当前界面缺乏实时语音反馈和可靠的用户打断机制。我们提出AsyncVoice Agent,其异步架构将流式LLM后端与对话式语音前端解耦,使叙述与推理并行运行,支持用户在任意时刻中断、提问或引导模型推理。客观测试表明,该方法相较单体基线将交互延迟降低超过600倍,同时保证高保真度与竞争性任务准确率。通过实现与模型思考过程的双向对话,AsyncVoice Agent为构建更高效、可操控、可信的人机系统提供了新范式,适用于医疗、金融等高风险任务。

原文摘要 · Abstract (English)

Effective human-AI collaboration on complex reasoning tasks requires that users understand and interact with the model's process, not just receive an output. However, the monolithic text from methods like Chain-of-Thought (CoT) prevents this, as current interfaces lack real-time verbalization and robust user barge-in. We present AsyncVoice Agent, a system whose asynchronous architecture decouples a streaming LLM backend from a conversational voice frontend. This design allows narration and inference to run in parallel, empowering users to interrupt, query, and steer the model's reasoning process at any time. Objective benchmarks show this approach reduces interaction latency by more than 600x compared to monolithic baselines while ensuring high fidelity and competitive task accuracy. By enabling a two-way dialogue with a model's thought process, AsyncVoice Agent offers a new paradigm for building more effective, steerable, and trustworthy human-AI systems for high-stakes tasks.

人机交互语音生成大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。