提出首个语音到语音翻译表达性评测基准,揭示现有系统在情感与非语言音素保留上的显著不足。
STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

- 构建32.6小时中英语音翻译表达性评测集,涵盖情感、语境风格与非语言音素等维度
- 六种系统中最佳情感保留仅3.82/5,非语言音素保留仅2.31/5,暴露表达性短板
- 采用基于大模型的属性对比框架,可有效模拟人类对表达性的判断
语音到语音翻译(S2ST)不仅需保持词汇意义准确,还需保留情感、语境风格(如新闻播报与戏剧对话)及非语言音素(NVs)。然而,大规模收集既忠实于原文又在表达上匹配源语音的跨语言目标语音极为困难,使基于参考的评估不可行。本文提出STEB(语音到语音翻译表达性基准),一个32.6小时的中英双语基准,用于评估标准维度(翻译忠实度、说话人相似度、时长对齐)与表达性维度(情感、语境风格、非语言音素保留)。针对表达性评估,STEB采用“描述-摘要”框架,将语音转换为结构化表达属性,并通过大模型判别器比较源与生成结果属性。人工验证显示,该方法在所有表达维度上与听者判断具有统计显著相关性。我们评估了六种S2ST系统,涵盖级联式、端到端模型及语音大模型。尽管多数系统在翻译忠实度上表现良好,但在情感保留(最佳:3.82/5)和非语言音素保留(最佳:2.31/5)方面仍存在明显缺陷。结果揭示了语义传递与表达传递之间的差距,明确指出表达性保留在当前S2ST中仍是一个开放挑战。音频样本见 https://cmots.github.io/steb.github.io/
原文摘要 · Abstract (English)
Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreover, collecting cross-lingual target speech that is both translation-faithful and expressively aligned with the source is difficult at scale, making reference-based evaluation impractical. We introduce STEB (Speech-to-Speech Translation Expressiveness Benchmark), a 32.6-hour Chinese--English benchmark that evaluates both standard dimensions (translation fidelity, speaker similarity, duration alignment) and expressiveness dimensions (emotion, scenario style, NV preservation). For expressiveness evaluation, STEB uses a caption-then-summarize framework that converts speech into structured expressive attributes and compares source and hypothesis attributes with an LLM judge. Human validation shows statistically significant correlations with listener judgments across all expressive dimensions. We evaluate six S2ST systems covering cascaded systems, end-to-end models, and speech large language models. Many systems, especially cascaded ones, achieve strong translation fidelity, but they still struggle with emotion preservation (best: 3.82/5) and NV preservation (best: 2.31/5). These results reveal a gap between semantic transfer and expressive transfer, identifying expressiveness preservation as an open challenge for S2ST. Audio samples are available at https://cmots.github.io/steb.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。