arXiv:2601.06329cs.CLcs.AI2026-01ACL

提出新评估方法,让语音模型评价更真实可靠。

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

  • 用似然与生成方法替代传统全局词元困惑度。
  • 新指标与人工评分相关性更强,能更好反映生成质量。
  • 发现顶尖模型与人类表现差距远小于旧评估结果。

基于大规模原始音频预训练的生成式语音语言模型可延续语音提示并保持说话人、情感等属性,是语音对话的基础模型。以往研究常使用「全局词元困惑度」评估这类模型,该方法直接将文本困惑度公式应用于语音标记。然而,这一做法忽略了语音与文本模态的根本差异,可能低估了语音特性。本文提出多种基于似然和生成的评估方法,替代传统的全局词元困惑度。实验表明,新评估方法更能准确反映生成质量,与人工评分的平均意见分数(MOS)相关性更强。在新指标下,语音模型的相对性能格局被重塑,最佳模型与人类基准之间的差距显著缩小。结果表明,恰当的评估对准确衡量语音建模进展至关重要。

原文摘要 · Abstract (English)

Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving attributes like speaker and emotion, serving as foundation models for spoken dialogue. In prior literature, these models are often evaluated using ``global token perplexity'', which directly applies the text perplexity formulation to speech tokens. However, this practice overlooks fundamental differences between speech and text modalities, possibly leading to an underestimation of the speech characteristics. In this work, we propose a variety of likelihood- and generative-based evaluation methods that serve in place of naive global token perplexity. We demonstrate that the proposed evaluations more faithfully reflect perceived generation quality, as evidenced by stronger correlations with human-rated mean opinion scores (MOS). When assessed under the new metrics, the relative performance landscape of spoken language models is reshaped, revealing a significantly reduced gap between the best-performing model and the human topline. Together, these results suggest that appropriate evaluation is critical for accurately assessing progress in spoken language modeling.

语音生成模型评估困惑度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。