arXiv:2507.12705cs.CLcs.SD2025-07Conference of the …被引 20

用大模型当语音评估裁判,统一判断发音、语速等质量。

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

  • 用音频拼接+上下文学习提升评估准确率
  • 多维度集成评估在基准上达0.91相关性
  • 适合语音系统评测与自动化测试场景

当前语音评估面临两大难题:需为不同音频特性定制专用系统,且自动评估与人类偏好相关性差。本文提出AudioJudge,系统研究大音频模型(LAM)作为评估裁判的可行性,覆盖发音、语速、说话人识别、语音质量等特征检测,以及系统级人类偏好模拟。通过对比不同提示工程策略,发现音频拼接结合上下文学习显著提升性能。进一步提出多方面集成的AudioJudge,将评估分解为词汇内容、语音质量与副语言特征三类专用裁判,在系统排名基准上实现最高0.91的斯皮尔曼相关性。鲁棒性分析显示,尽管在噪声环境下表现稳定,但大模型存在明显冗余和位置偏差,需谨慎处理。

原文摘要 · Abstract (English)

Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation methods and human preferences. This work presents a systematic study of Large Audio Model (LAM) as a Judge, AudioJudge, investigating whether it can provide a unified evaluation framework that addresses both challenges. We systematically explore AudioJudge across audio characteristic detection tasks, including pronunciation, speaking rate, speaker identification and speech quality, and system-level human preference simulation for automated benchmarking. We investigate different prompt engineering strategies, finding that audio concatenation combined with in-context learning significantly improves performance across both audio characteristic detection and human preference simulation tasks. We further introduce a multi-aspect ensemble AudioJudge to enable general-purpose multi-aspect audio evaluation. This method decomposes speech assessment into specialized judges for lexical content, speech quality, and paralinguistic features, achieving up to 0.91 Spearman correlation with human preferences on our system ranking benchmark. Robustness analysis reveals that while LAMs maintain strong performance under acoustic noise, they exhibit significant verbosity and positional biases that require careful mitigation.

语音评估大模型自动评测多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。