提出双轴生成式奖励模型,提升对话系统语义与对话节奏的评估精度。
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models

- 构建双轴奖励模型,分别评估语义质量和对话时机。
- 在多类数据集上达到当前最佳对话质量评估效果。
- 适合需要精细反馈的对话系统强化学习研究者使用。
实现无缝、类人交互仍是全双工语音对话模型(SDMs)的关键挑战。强化学习(RL)已显著提升文本与视觉-语言模型性能,而精心设计的奖励信号对RL表现至关重要。我们认为RL是解决SDM核心挑战的有前景策略。然而,一个根本性障碍依然存在:现有自动化评估指标依赖行为统计或时间预测准确率等表面代理,无法为RL提供可靠奖励信号。另一方面,人工评估虽丰富,但成本高、不一致且难以扩展。为此,我们提出双轴生成式奖励模型,通过详细分类体系和标注数据集训练,理解复杂的交互动态,输出单一评分及关键的语义质量与交互时机分项评估。双重输出为SDMs提供精准诊断反馈,并生成适用于在线强化学习的可靠、可指导的奖励信号。该模型在涵盖合成对话与复杂真实交互的多种数据集上均达到对话质量评估的最先进水平。
原文摘要 · Abstract (English)
Achieving seamless, human-like interaction remains a key challenge for full-duplex spoken dialogue models (SDMs). Reinforcement learning (RL) has substantially enhanced text- and vision-language models, while well-designed reward signals are crucial for the performance of RL. We consider RL a promising strategy to address the key challenge for SDMs. However, a fundamental barrier persists: prevailing automated metrics for assessing interaction quality rely on superficial proxies, such as behavioral statistics or timing-prediction accuracy, failing to provide reliable reward signals for RL. On the other hand, human evaluations, despite their richness, remain costly, inconsistent, and difficult to scale. We tackle this critical barrier by proposing a Dual-Axis Generative Reward Model, which is trained to understand complex interaction dynamics using a detailed taxonomy and an annotated dataset, produces a single score and, crucially, provides separate evaluations for semantic quality and interaction timing. Such dual outputs furnish precise diagnostic feedback for SDMs and deliver a dependable, instructive reward signal suitable for online reinforcement learning. Our model achieves state-of-the-art performance on interaction-quality assessment across a wide spectrum of datasets, spanning synthetic dialogues and complex real-world interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。