arXiv:2606.31055cs.CLcs.SD2026-06被引 1

为对话系统语音韵律提供可解释的评估方法。

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

  • 基于真实对话数据构建匹配的参考基准,精准衡量语音特征。
  • 相比传统方法,误报率更接近理论值10%,结果更可信。
  • 适合语音合成与对话系统开发者用于优化自然度。

语音到语音(S2S)AI代理发展迅速,但缺乏可解释的语音原生评价指标来衡量对话中的语调和节奏。由于基频(F₀)、语速、发音速率和停顿模式会随模型预测的说话人特征和交互状态变化,使用汇总的人类统计数据进行评估可能不准确。基于来自Seamless Interaction数据集的4000+小时双人英语对话,我们构建了匹配的参考区间,涵盖F₀均值、F₀表现力、语速、发音速率、停顿比例和平均停顿时长。我们提出一种百分位评估协议:从S2S输出波形中提取相同指标,与最接近的匹配人类参考分组比较,报告百分位偏差或5th-95th百分位外的异常标记。在保留的人类样本上,汇总参考过度标记状态相关的F₀表现力和节奏,而匹配参考则使标记率接近名义上的10%,且偏差方向具有可解释性。这些输出作为行为合理性检查,补充而非替代感知和以用户为中心的评估。

原文摘要 · Abstract (English)

Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because $F_0$, speaking rate, articulation rate, and pausing shift with model-predicted speaker traits and interaction state, pooled human statistics can be poorly calibrated for evaluating a particular output. Using 4000+ hours of dyadic English conversation from the Seamless Interaction dataset, we construct matched reference regimes for $F_0$ mean, $F_0$ expressivity, speech rate, articulation rate, pause ratio, and mean pause duration. We then define a percentile-based evaluation protocol: extract the same metrics from an S2S output waveform, compare them to the closest matched human reference stratum, and report percentile deviations or 5th-95th percentile out-of-regime flags. On held-out human rows, pooled references over-flag state-conditioned $F_0$ expressivity and rhythm, while matched references return flag rates closer to the nominal 10% and make deviation direction interpretable. These outputs serve as behavioral plausibility checks that complement, rather than replace, perceptual and user-centered evaluation.

语音评估对话系统韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。