arXiv:2506.21463cs.CLcs.LG2025-06ICML被引 16

用用户对话数据训练语音对话模型,提升实时交互的自然性与安全性。

Aligning Spoken Dialogue Models from User Interactions

  • 基于真实多轮语音对话构建超15万条偏好对数据集
  • 通过离线微调使模型在事实性、安全性上显著提升
  • 适用于需要自然实时语音交互的场景,如智能客服

我们提出一种新型偏好对齐框架,用于提升语音对话模型在真实实时对话中的表现。现有偏好学习方法主要针对文本语言模型,难以适配语音交互中丰富的动态特征(如打断、插话)及无明确发言轮次分割的问题。为此,我们从原始多轮语音对话中构建了超过15万条带人工智能反馈的偏好对齐样本,覆盖语言内容与时间上下文变化。采用离线对齐方法微调全双工自回归语音到语音模型。大量实验表明,对通用对话的反馈可持续提升模型生成更准确、更安全、更符合上下文的交互内容。我们部署微调后模型并开展全面人工评估,验证其在非单轮对话中的综合效果。研究揭示了各类动态因素间良好平衡的重要性,这对构建自然实时语音对话系统至关重要。

原文摘要 · Abstract (English)

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not directly suited to the complexities of real-time speech interactions, with richer dynamics (e.g. interruption, interjection) and no explicit segmentation between speaker turns.We create a large-scale dataset of more than 150,000 preference pairs from raw multi-turn speech conversations, annotated with AI feedback, to cover preferences over both linguistic content and temporal context variations. We leverage offline alignment methods to finetune a full-duplex autoregressive speech-to-speech model. Extensive experiments demonstrate that feedback on generic conversations can be consistently effective in improving spoken dialogue models to produce more factual, safer and more contextually aligned interactions. We deploy the finetuned model and conduct holistic human evaluations to assess the impact beyond single-turn conversations. Our findings shed light on the importance of a well-calibrated balance among various dynamics, crucial for natural real-time speech dialogue systems.

语音对话偏好对齐实时交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。