实时对话系统中,帧级同步预测说话人切换状态。
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

- 共享流表示下并行处理语音识别与说话人切换状态
- 在中英双语数据集上实现高精度低延迟的判断
- 适合需要快速响应的实时对话系统应用
准确且及时的说话人切换对语音对话系统至关重要,必须实时区分用户打断、可忽略的应答词以及话语结束。以往模块化方法通常在话语或固定片段级别优化切换状态预测,与连续状态估计不匹配,且常依赖额外的语音识别模型,限制响应速度并增加系统复杂度。为此,我们提出 X2-Turn,一种基于延迟流建模的帧同步切换状态预测方法。具体而言,在预训练的 Voxtral Realtime 模型基础上,引入一个与语音识别头并行的帧同步切换状态头,共同在共享流表示上进行推理,实现在帧级别联合预测语音识别结果与细粒度说话人切换状态。我们在双语中英文 Easy-Turn 测试集上评估该方法,结果表明其在保持低延迟的同时实现了精确的说话人切换检测。
原文摘要 · Abstract (English)
Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level, creating a mismatch with the continuous turn state estimate, and often depend on an auxiliary ASR model, which limits responsiveness and increases overall system complexity. Therefore, we present X2-Turn, a frame-synchronous turn state prediction method via delayed-stream modeling. Specifically, building on the pretrained Voxtral Realtime model, we introduce a frame-synchronous turn state head that operates in parallel with the ASR head on shared streaming representations, jointly predicting ASR tokens and fine-grained turn states at the frame level. We evaluate our method on the bilingual Chinese-English Easy-Turn test sets, and the results demonstrate its effectiveness in achieving accurate turn-taking detection while maintaining low latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。