arXiv:2410.16011cs.CLcs.AI2024-10NAACL被引 2

纠正语音翻译延迟评估的误区,提出更准确的计算感知延迟测量方法

CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation

  • 发现现有延迟评估方法因基础误解导致结果偏高
  • 验证该问题在流式与分段场景中均普遍存在
  • 提出修正方案,使延迟测量更贴近真实系统表现

同时性语音翻译(SimulST)系统需在翻译质量与响应速度间取得平衡,延迟测量对评估其实际性能至关重要。然而,长期以来存在一种观点认为当前指标在非分段流式设置下会给出过高的延迟值。本文对此现象展开研究,揭示其根源在于现有延迟评估方法中的根本性误解。我们证明该问题不仅影响流式场景,也存在于不同指标下的分段级延迟评估中。此外,本文提出一种改进方法,可正确测量SimulST系统的计算感知延迟,解决现有指标的局限性。

原文摘要 · Abstract (English)

Simultaneous speech translation (SimulST) systems must balance translation quality with response time, making latency measurement crucial for evaluating their real-world performance. However, there has been a longstanding belief that current metrics yield unrealistically high latency measurements in unsegmented streaming settings. In this paper, we investigate this phenomenon, revealing its root cause in a fundamental misconception underlying existing latency evaluation approaches. We demonstrate that this issue affects not only streaming but also segment-level latency evaluation across different metrics. Furthermore, we propose a modification to correctly measure computation-aware latency for SimulST systems, addressing the limitations present in existing metrics.

语音翻译延迟评估模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。