arXiv:2512.17648cs.CL2025-12中稿 · EMNLP被引 6

首个开源工具箱,统一评估流式语音翻译系统性能与交互演示。

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

  • 支持连续长语音的增量与重译解码,覆盖真实应用场景。
  • 提供细粒度日志,可量化评估翻译质量与延迟表现。
  • 内置交互式网页界面,实现实时可视化对比与演示。

流式语音转文本翻译(StreamST)需在严格延迟约束下,随语音输入同步生成翻译结果,要求模型兼顾低延迟与高翻译质量。尽管进展迅速,现有评估框架仍碎片化严重,对系统运行假设不一——例如是否处理连续语音或短片段音频、是否支持输出修正(重译)等。以SimulEval为例,其仅支持增量解码、假设短段输入,且缺乏原生系统演示功能。导致跨研究公平比较困难,缺乏统一基准与交互展示方案。为此,我们提出simulstream,首个专为StreamST设计的开源评估与演示工具箱,支持长语音下的增量与重译解码,提供细粒度日志用于质量与延迟评估,并集成交互式网页界面,实现实时可视化与系统对比。

原文摘要 · Abstract (English)

Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality. Despite rapid progress, evaluation remains fragmented across existing frameworks, which make different assumptions about how systems operate - for example, whether they process continuous speech or short pre-segmented audio, and whether they support output revision (retranslation) or not (incremental). For instance, SimulEval, the most widely used framework, supports only incremental decoding, assumes short segmented inputs, and lacks a native support for system demonstrations. As a result, comparing systems fairly and consistently across studies remains challenging, with no unified solution for benchmarking and interactive demonstration. To address this gap, we introduce simulstream, the first open-source framework for StreamST evaluation and demonstration. It supports both incremental and re-translation decoding on long-form speech, provides fine-grained logging for quality and latency evaluation, and includes an interactive web interface for real-time visualization and comparison.

流式翻译语音识别开源工具评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。