arXiv:2604.16456cs.CLcs.AI2026-04被引 2

评测语音助手在被打断后修正任务状态的能力,发现现有模型表现普遍不佳。

EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions

  • 构建可控的全双工中断场景,标准化测试语音助手的状态更新能力。
  • 40.2%错误源于打断后的状态推理,而非任务本身难度。
  • 所有测试模型在中断后任务通过率均未超50%,适合改进语音交互系统研发者。

实时语音助手需在用户中途打断时修正任务状态,但现有对话基准多评估轮次式交互,忽略此关键问题。本文提出EchoChain,一个针对全双工状态下中断时状态更新推理的可控基准。该基准识别出三种常见失败模式:上下文惯性、中断失忆和目标偏离。通过生成情景驱动的对话,并在助理发言起始后标准化位置注入打断,实现跨模型的可控对比。在配对半双工对照实验中,总错误率比中断运行降低40.2%,表明多数错误源自中断下的状态推理,而非任务本身难度。在评估的实时语音模型中,无一系统通过率超过50%,凸显了当前模型在生成过程中状态修订能力的巨大提升空间。EchoChain为诊断全双工语音交互中的状态更新失败提供了可复现的评测框架。

原文摘要 · Abstract (English)

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a controlled benchmark for evaluating full-duplex state-update reasoning under mid-speech interruptions. EchoChain identifies three recurring failure patterns in post-interruption continuations: contextual inertia, interruption amnesia, and objective displacement. The benchmark generates scenario-driven conversations and injects interruptions at a standardized point relative to assistant speech onset, enabling controlled cross-model comparison. In a paired half-duplex control, total failures drop by 40.2% relative to interrupted runs, indicating that many errors are driven by state-update reasoning under interruption rather than task difficulty alone. Across evaluated real-time voice models, no system exceeds a 50% pass rate, showing substantial room for improvement in mid-generation state revision. EchoChain provides a reproducible benchmark for diagnosing state-update reasoning failures in full-duplex voice interaction.

语音助手状态更新全双工对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。