arXiv:2606.24648cs.SDcs.CL2026-06中稿 · Interspeech 2026被引 1

构建语音情感维度评估基准,揭示大模型判别能力不足。

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

论文配图:ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
图 1 · 摘自论文原文
  • 设计5175对音频的配对评测集,覆盖5类语音情感特征。
  • 大模型平均比人类差32个百分点,且在模糊判断时错误率高。
  • 支持跨文本与同文本测试,可分析模型依赖语音或语义的倾向。

大型音频语言模型(LALM)已被广泛用作生成语音自动评估的评判模型。然而,现有方法主要关注整体自然度,对细微的副语言特征区分仍缺乏探索。我们提出 ParaPairAudioBench,一个包含5,175组音频对的成对评测基准,涵盖风格、语速、强调、年龄和性别五个副语言维度。实验表明,当前LALM评判模型平均比人类低32个百分点,尤其在应放弃判断的平局情况中存在严重校准失败。为分析词汇与声学特征的依赖程度,该基准包含同转录与跨转录两种条件。ParaPairAudioBench 支持多维度、校准感知的评估,可用于检验 LALM 作为评判者在副语言语音评估中的可靠性。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored. We introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. Our experiments show that current LALM judges still lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases where the correct decision is to abstain. To further analyze lexical versus acoustic reliance, the benchmark includes both same-transcript and cross-transcript conditions. ParaPairAudioBench enables multi-dimensional, calibration-aware assessment of the reliability of LALM-as-a-Judge for paralinguistic speech evaluation.

语音评估大模型评测副语言校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。