用多维度反馈优化语音对话系统,让回复更自然连贯。
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
- 设计多奖励强化学习框架,同时优化语义、音质和情感一致性。
- 联合训练使语义质量和音频自然度同步提升,优于单一目标优化。
- 适用于需要实时响应的语音对话系统,适合语音交互研发者。
针对语音输入/输出对话系统(SDS),现有基于人类或人工智能反馈的强化学习(RLHF/RLAIF)研究仍不充分,多数工作仅在话语层面使用单一语义奖励。此类方法忽视了对话质量的多维性与多模态特性,包括语义连贯性、语音自然度、说话人一致性、情感对齐及对话轮次行为。此外,它们与双工语音系统中逐块增量生成的机制不匹配。本文提出首个面向SDS的多奖励RLAIF框架,融合语义、音频质量与情感一致性奖励。为对齐话语级偏好与双工模型的增量式、分块解码过程,采用轮次级偏好采样,并在单个直接策略优化(DPO)目标中聚合每块的日志概率。我们首次系统性地研究了在多轮思维链与分块双工模型中通过偏好学习提升对话质量的方法,并发布了一个多奖励DPO数据集以支持可复现研究。实验表明,单奖励RLAIF仅提升对应指标,而联合多奖励训练在语义质量与音频自然度上均取得一致改进,凸显了整体化多奖励对实际对话系统的必要性。
原文摘要 · Abstract (English)
Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single semantic rewards applied at the utterance level. Such setups overlook the multi-dimensional and multi-modal nature of conversational quality, which encompasses semantic coherence, audio naturalness, speaker consistency, emotion alignment, and turn-taking behavior. Moreover, they are fundamentally mismatched with duplex spoken dialogue systems that generate responses incrementally, where agents must make decisions based on partial utterances. We address these limitations with the first multi-reward RLAIF framework for SDS, combining semantic, audio-quality, and emotion-consistency rewards. To align utterance-level preferences with incremental, blockwise decoding in duplex models, we apply turn-level preference sampling and aggregate per-block log-probabilities within a single DPO objective. We present the first systematic study of preference learning for improving SDS quality in both multi-turn Chain-of-Thought and blockwise duplex models, and release a multi-reward DPO dataset to support reproducible research. Experiments show that single-reward RLAIF selectively improves its targeted metric, while joint multi-reward training yields consistent gains across semantic quality and audio naturalness. These results highlight the importance of holistic, multi-reward alignment for practical conversational SDS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。