厘清全双工对话系统的核心差异,提出三套分类框架
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

- 按决策层级划分架构,明确全双工位置
- 定义交互类型与响应关系,统一术语标准
- 提供状态机模型,适合系统设计者参考
近年来多个语音对话系统自称具备“全双工”能力,但实际含义差异巨大。现有综述将其简单归为端到端/级联或工程化/学习型,忽略了对构建者至关重要的区别。本文指出核心问题在于术语缺乏清晰的分类体系:当前定义未说明全双工决策发生在何处、支持哪些交互类型、以及系统如何实时行为。为此提出三个互补框架:(i) L0-L3 架构层级,定位全双工决策位置;(ii) T×I×R 交互本体,明确每类交互的时序关系、用户意图和系统响应要求;(iii) 决策状态机(IDLE/LISTEN/SPEAK/WAIT/DUAL),描述系统状态流转。通过审计已有系统与基准,发现实现差距显著:尽管许多架构理论上可支持全双工,其实际行为仍受限于训练与评估中所覆盖的交互模式。我们指出,公开数据集覆盖有限,工业级语料大多未披露,且尚未实现L3表示层建模,是未来研究的关键前沿。相关材料见 https://github.com/DuplexLM/DuplexSurvey。
原文摘要 · Abstract (English)
More than a dozen spoken dialogue systems have recently claimed to be "full-duplex," yet the term has been used to describe substantially different capabilities. Existing surveys collapse them onto a single axis (cascaded/end-to-end, or engineered/learned) and miss the distinctions that matter most for builders. We argue that much of this ambiguity is taxonomical: current terminology does not specify where duplex decisions are made, which interaction types are supported, or how a system behaves moment by moment. This paper introduces three complementary frameworks: (i) an L0-L3 Architectural Hierarchy that locates where duplex decisions are made; (ii) a $T\times I\times R$ Interaction Ontology that specifies the temporal relation, user intent, and required system response for each interaction; and (iii) a Decision State Machine (IDLE/LISTEN/SPEAK/WAIT/DUAL) that describes how systems move between states. Across published systems and benchmarks, our audit documents a realization gap: although many architectures can in principle operate in full-duplex states, their observed behavior remains constrained by the interaction patterns represented in training and evaluation. We point to the limited public training-data coverage relative to the (largely undisclosed) industrial corpora, together with the still-unrealized goal of L3 representation-level modeling, as the key frontiers for future research on full-duplex dialogue. The related material is available at https://github.com/DuplexLM/DuplexSurvey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。