arXiv:2604.20842cs.CLcs.AI2026-04被引 2

构建首个覆盖超100种语音特征的语音生成评测基准

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation

论文配图:SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
图 1 · 摘自论文原文
  • 设计三阶段渐进任务评估语音情感等细粒度控制能力
  • 通过大模型对比打分实现无人工标注的稳定评估
  • 发现主流模型43.3%错误源于对语气线索误读

语音韵律特征对自然人机交互至关重要,但当前大型音频语言模型在该领域的评估仍受限于粗粒度特征覆盖与主观判断。为此,我们提出SpeechParaling-Bench,一个全面的韵律感知语音生成评测基准。其将特征覆盖从不足50项扩展至超过100项,包含1,000多对中英平行语音查询,并分为三个逐级递增难度的任务:细粒度控制、句内变化与情境自适应。为实现可靠评估,我们开发了一套基于大模型判官的成对比较流程,通过相对偏好判断替代绝对评分,有效降低主观性,避免昂贵的人工标注,同时提升评估稳定性与可扩展性。大量实验表明当前主流大模型存在显著缺陷:即使领先专有模型也难以实现完整的静态控制与动态调制,且在情境对话中43.3%的错误源于未能正确理解韵律线索,凸显了迈向更鲁棒韵律建模以实现人类对齐语音助手的迫切需求。

原文摘要 · Abstract (English)

Paralinguistic cues are essential for natural human-computer interaction, yet their evaluation in Large Audio-Language Models (LALMs) remains limited by coarse feature coverage and the inherent subjectivity of assessment. To address these challenges, we introduce SpeechParaling-Bench, a comprehensive benchmark for paralinguistic-aware speech generation. It expands existing coverage from fewer than 50 to over 100 fine-grained features, supported by more than 1,000 English-Chinese parallel speech queries, and is organized into three progressively challenging tasks: fine-grained control, intra-utterance variation, and context-aware adaptation. To enable reliable evaluation, we further develop a pairwise comparison pipeline, in which candidate responses are evaluated against a fixed baseline by an LALM-based judge. By framing evaluation as relative preference rather than absolute scoring, this approach mitigates subjectivity and yields more stable and scalable assessments without costly human annotation. Extensive experiments reveal substantial limitations in current LALMs. Even leading proprietary models struggle with comprehensive static control and dynamic modulation of paralinguistic features, while failure to correctly interpret paralinguistic cues accounts for 43.3% of errors in situational dialogue. These findings underscore the need for more robust paralinguistic modeling toward human-aligned voice assistants.

语音生成韵律建模评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。