首个评估人机全双工语音交互的基准,重点测试紧急情况下的打断能力。
FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
- 构建六种真实场景,测试模型在对话中的响应延迟与中断能力。
- 开源模型在紧急识别和快速终止对话上显著落后于商业模型。
- 提出新预测方式,有望实现更自然流畅的实时对话体验。
全双工语音到语音大语言模型是实现自然人机交互的基础,支持实时语音对话系统。然而,对这类模型的评测与建模仍是重大挑战。我们提出了FLEXI,首个专门针对全双工人机语音交互的基准,明确引入紧急情境下的模型打断机制。FLEXI通过六种多样化的真人-模型交互场景,系统评估了实时对话中的延迟、语音质量与对话有效性,揭示了开源模型与商用模型在紧急感知、话轮终止和交互延迟方面存在显著差距。最后,我们提出下一刻词对预测可能是实现真正无缝、类人全双工交互的可行路径。
原文摘要 · Abstract (English)
Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fundamental challenge. We introduce FLEXI, the first benchmark for full-duplex LLM-human spoken interaction that explicitly incorporates model interruption in emergency scenarios. FLEXI systematically evaluates the latency, quality, and conversational effectiveness of real-time dialogue through six diverse human-LLM interaction scenarios, revealing significant gaps between open source and commercial models in emergency awareness, turn terminating, and interaction latency. Finally, we suggest that next token-pair prediction offers a promising path toward achieving truly seamless and human-like full-duplex interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。