arXiv:2509.17349cs.CLcs.AI2025-09被引 5

提出新延迟评估方法,让实时语音翻译更准。

Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation

  • 设计新指标YAAL和LongYAAL,解决分割带来的偏差
  • 在短时与长时场景下均显著优于现有指标
  • 配套重分段工具SoftSegmenter,适合开发者部署

同时性语音转文本翻译需权衡译文质量与延迟。尽管质量评估已成熟,延迟测量仍存挑战,现有指标在短时场景(人工预分割)下结果不一致。本文首次对跨语言对与系统组合的延迟指标进行综合元评估,发现当前指标存在与分割相关的结构性偏差。提出YAAL(Yet Another Average Lagging)用于短时评估,LongYAAL用于无分割音频。设计SoftSegmenter,基于软词级对齐实现重分段。实验表明,结合YAAL、LongYAAL与SoftSegmenter可更可靠评估短时与长时同时性语音翻译系统。所有工具均已集成于OmniSTEval框架:https://github.com/pe-trik/OmniSTEval。

原文摘要 · Abstract (English)

Simultaneous speech-to-text translation systems must balance translation quality with latency. Although quality evaluation is well established, latency measurement remains a challenge. Existing metrics produce inconsistent results, especially in short-form settings with artificial presegmentation. We present the first comprehensive meta-evaluation of latency metrics across language pairs and systems. We uncover a structural bias in current metrics related to segmentation. We introduce YAAL (Yet Another Average Lagging) for a more accurate short-form evaluation and LongYAAL for unsegmented audio. We propose SoftSegmenter, a resegmentation tool based on soft word-level alignment. We show that YAAL and LongYAAL, together with SoftSegmenter, outperform popular latency metrics, enabling more reliable assessments of short- and long-form simultaneous speech translation systems. We implement all artifacts within the OmniSTEval toolkit: https://github.com/pe-trik/OmniSTEval.

语音翻译延迟评估AI测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。