arXiv:2605.28618eess.AS2026-05ACL被引 3

构建语音生成评测基准,揭示模型在长文本一致性上的不足

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

论文配图:Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
图 1 · 摘自论文原文
  • 拆解长语音质量为声学、语义、表现力等独立维度进行评估
  • 覆盖17种场景1101个样本,设计7项自动化指标
  • 发现当前模型在高表现力场景下仍存明显差距

语音合成技术虽已实现高保真度,但针对长上下文条件下的系统性评估仍不充分。现有测试场景多局限于特定领域,与多样化下游应用存在显著差距;同时,现有评价指标忽视了连贯性与一致性等关键长文本因素,难以可靠泛化。为此,我们提出Swanbench-Speech,一个将长语音质量分解为具体、解耦维度的综合性评测基准。该基准具有三大特性:1)丰富语音场景:聚焦长语音生成与对话生成,涵盖声学、语义与表现力挑战,包含1,101个样本,覆盖17类常见语音场景;2)全面评估维度:沿声学、语义与表现力轴线,定义七项自动化评估指标,实现全面、准确、标准化评估;3)重要洞见:通过大规模实验发现,当前模型在高度表现力场景中仍表现不佳,且在连贯性与层次结构上显著落后于真实录音。

原文摘要 · Abstract (English)

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A comprehensive evaluation benchmark for long-form speech is indispensable for two reasons: 1) existing test scenarios are often confined to limited domains, creating a significant gap with the diverse downstream applications; 2) existing metrics overlook critical long-text factors such as consistency and coherence, failing to generalize reliably. To this end, we propose Swanbench-Speech, a comprehensive benchmark that decomposes long-form speech quality into specific, disentangled dimensions. SwanBench-Speech has three key properties. 1) Rich speech scenarios: Focusing on long-form speech generation and dialog generation, SwanBench-Speech covers acoustics, semantics, and expressiveness challenges, and consists of 1,101 samples spanning 17 common speech scenarios; 2) Comprehensive evaluation dimensions: Along the acoustics, semantics, and expressiveness axes, SwanBench-Speech defines an automated evaluation protocol with seven metrics to provide a comprehensive, accurate, and standardized assessment; 3) Valuable Insights: Through extensive experiments, we reveal that current models still struggle in highly expressive scenarios and exhibit a notable gap in consistency and hierarchy compared to real recordings.

语音生成评测基准长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。