arXiv:2509.14515cs.CLcs.SD2025-09综述被引 10

让AI实现真正同步对话,像人一样边听边说、随时插话。

From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models

  • 分两类架构:模块化设计和端到端学习,实现说话与倾听同步。
  • 现有模型在同步数据、结构设计和评测标准上存在明显短板。
  • 适合研究人机交互、语音模型和对话系统的人参考。

真正的全双工(TFD)语音通信——支持自然的轮换对话、重叠说话和打断——是迈向类人智能交互的关键里程碑。本文全面综述了大语言模型时代下的全双工语音语言模型(FD-SLMs)。我们建立了一个分类体系,区分工程化同步(模块化架构)与学习型同步(端到端架构),并将碎片化的评估方法统一为涵盖时间动态、行为仲裁、语义连贯性和声学性能的框架。通过对主流FD-SLMs的对比分析,识别出根本挑战:同步数据稀缺、架构差异大、评测不一致,并为推进人机语音交互提供路线图。

原文摘要 · Abstract (English)

True Full-Duplex (TFD) voice communication--enabling simultaneous listening and speaking with natural turn-taking, overlapping speech, and interruptions--represents a critical milestone toward human-like AI interaction. This survey comprehensively reviews Full-Duplex Spoken Language Models (FD-SLMs) in the LLM era. We establish a taxonomy distinguishing Engineered Synchronization (modular architectures) from Learned Synchronization (end-to-end architectures), and unify fragmented evaluation approaches into a framework encompassing Temporal Dynamics, Behavioral Arbitration, Semantic Coherence, and Acoustic Performance. Through comparative analysis of mainstream FD-SLMs, we identify fundamental challenges: synchronous data scarcity, architectural divergence, and evaluation gaps, providing a roadmap for advancing human-AI communication.

全双工语音交互对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。