arXiv:2603.14889eess.AScs.CL2026-03ACL被引 1

构建对话奖励模型,同时评估语音情感与口语自然度。

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

  • 直接处理多轮语音对话,联合建模语调与口语化特征。
  • 在对话偏好任务中准确率超越通用音频大模型。
  • 适合研究语音对话系统评价与自然语言生成优化者。

端到端语音对话系统的快速发展要求超越单纯的文本语义,融入语音中的副语言细节和人类对话的自发性。然而,现有方法存在两大关键缺陷:语调与情感等声学特质的模态鸿沟,以及书面语与自然口语之间的口语化鸿沟。为此,我们提出SDiaReward,一个基于全新构建的SDiaReward-Dataset(episode-level preference pairs)训练的端到端多轮对话奖励模型。该模型直接处理完整多轮语音对话,通过成对偏好监督进行优化,可在一个评估器中联合衡量模态与口语化程度。我们进一步构建了ESDR-Bench,一个分层化的基准测试集,用于鲁棒的端到端评估。实验表明,SDiaReward在成对偏好判断上达到当前最优性能,显著优于通用音频大模型。进一步分析表明,其能捕捉对话表达力的相对差异,而非仅依赖表面合成特征,从而提升跨领域与录音条件的泛化能力。代码、数据与演示已公开于https://github.com/MM-Speech/SDiaReward/。

原文摘要 · Abstract (English)

The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances and the spontaneous nature of human conversation. However, current methods struggle with two critical gaps: the modality gap, involving prosody and emotion, and the colloquialness gap, distinguishing written scripts from natural speech. To address these challenges, we introduce SDiaReward, an end-to-end multi-turn reward model trained on SDiaReward-Dataset, a novel collection of episode-level preference pairs explicitly targeting these gaps. It operates directly on full multi-turn speech episodes and is optimized with pairwise preference supervision, enabling joint assessment of modality and colloquialness in a single evaluator. We further establish ESDR-Bench, a stratified benchmark for robust episode-level evaluation. Experiments demonstrate that SDiaReward achieves state-of-the-art pairwise preference accuracy, significantly outperforming general-purpose audio LLMs. Further analysis suggests that SDiaReward captures relative conversational expressiveness beyond superficial synthesis cues, improving generalization across domains and recording conditions. Code, data, and demos are available at https://github.com/MM-Speech/SDiaReward/.

语音对话奖励模型口语化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。