arXiv:2606.03241cs.CLeess.AS2026-06被引 1

构建统一评估框架,让语音翻译模型比拼更公平可靠。

Benchmarking Speech-to-Speech Translation Models

论文配图:Benchmarking Speech-to-Speech Translation Models
图 1 · 摘自论文原文
  • 设计涵盖46项指标的统一评测框架,覆盖8个维度。
  • 发现不同架构在自然度上差距超30%,单一指标易误导评价。
  • 筛选出10项关键指标,节省2.5倍评估时间且保留人类判断相关性。

语音到语音翻译(S2ST)发展迅速,但离线评估缺乏统一标准:研究使用互不重叠的指标集,难以直接比较。本文提出COMPASS,一个统一且可复现的基准评测框架,整合46项指标,覆盖八个维度,并在来自FLEURS和CVSS的1,248个模型-语言配置上部署,涵盖十种语言对的级联与端到端架构。结果显示,不同架构在自然度和说话人保真度上的最佳与最差差距超过30%,而翻译质量差异仅几百分点,因此单指标排名会系统性扭曲模型真实表现。通过相关性过滤,将46项指标压缩为每方向10项,其中X→EN与EN→X需不同指标组合(如TER/UTMOS vs. ChrF++/NISQA-MOS),该子集保持排名一致性(斯皮尔曼ρ > 0.80),同时评估时间减少约2.5倍。跨配音、播客、医疗领域的真人验证表明,独立的MOS预测器无法准确反映听者偏好,而顶级领域特定指标与人类判断高度相关(ρ ≥ 0.90)。COMPASS已开源,作为面向领域的S2ST评估基础。

原文摘要 · Abstract (English)

Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets, preventing direct comparisons. We introduce COMPASS, a unified and reproducible benchmarking framework integrating 46 metrics across eight dimensions, and deploy it on 1,248 model-language configurations from FLEURS and CVSS, spanning cascaded and end-to-end architectures over ten language pairs. Architectures exhibit complementary strengths: best-vs-worst gaps exceed 30\% on naturalness and speaker preservation but remain within a few points on translation quality, so single-metric rankings systematically misrepresent system quality. Correlation filtering reduces 46 metrics to 10 per direction, with three axes requiring different metrics across X$\to$EN and EN$\to$X (e.g., TER/UTMOS vs. ChrF++/NISQA-MOS); these subsets preserve rankings (Spearman's $ρ>0.80$) while cutting evaluation time by $\approx 2.5\times$. Human validation across dubbing, podcasts, and medical domains shows standalone MOS predictors fail to predict listener preference, while top domain-specific metrics correlate with human judgment ($ρ\geq 0.90$). We release COMPASS as a foundation for domain-aware S2ST evaluation.

语音翻译评估框架多维指标可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。