让语音翻译系统像人类译员一样实时判断何时输出,提升效率与质量。
SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
- 模仿人类译员,感知新语义单元即触发翻译输出。
- 在延迟与翻译质量权衡上优于现有方法,决策速度提升9.6倍。
- 无需复杂训练数据和大模型推理,适合实时场景部署。
如何让同步语音翻译系统做出类人类译员的读写决策?当前最先进系统将同步语音翻译建模为多轮对话任务,需要专门的交错训练数据,并依赖计算开销大的大语言模型进行决策。本文提出SimulSense框架,通过持续接收输入语音并感知新语义单元来触发翻译输出,模拟人类译员行为。实验对比两种先进基线系统表明,所提方法在质量-延迟权衡上表现更优,且决策速度比基线快达9.6倍,显著提升实时效率。
原文摘要 · Abstract (English)
How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn dialogue task, requiring specialized interleaved training data and relying on computationally expensive large language model (LLM) inference for decision-making. In this paper, we propose SimulSense, a novel framework for SimulST that mimics human interpreters by continuously reading input speech and triggering write decisions to produce translation when a new sense unit is perceived. Experiments against two state-of-the-art baseline systems demonstrate that our proposed method achieves a superior quality-latency tradeoff and substantially improved real-time efficiency, where its decision-making is up to 9.6x faster than the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。