厘清实时语音翻译系统的真实挑战,推动研究向真实场景靠拢。
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
- 定义实时语音翻译系统的关键步骤与组件,建立统一术语体系。
- 分析110篇论文发现,现有研究多基于人工分割语音,脱离真实场景。
- 提出评估框架与架构改进建议,助力更贴近实际应用的系统发展。
同时语音转文本翻译(SimulST)在说话人发声的同时进行源语言语音到目标语言文本的翻译,以实现低延迟,提升用户理解。尽管其旨在处理无限长语音,但多数研究仍聚焦于人工预分割的语音,简化了任务,忽略了关键挑战。这种局限性,加上术语使用混乱,严重限制了研究成果在真实场景中的应用,阻碍了领域进展。我们对110篇论文进行广泛文献综述,不仅揭示了当前研究中的核心问题,还成为主要贡献的基础:1)定义SimulST系统的步骤与核心组件,提出标准化术语与分类体系;2)深入分析社区研究趋势;3)提出具体建议与未来方向,涵盖评估框架到系统架构,推动领域向更真实、高效的SimulST解决方案演进。
原文摘要 · Abstract (English)
Simultaneous speech-to-text translation (SimulST) translates source-language speech into target-language text concurrently with the speaker's speech, ensuring low latency for better user comprehension. Despite its intended application to unbounded speech, most research has focused on human pre-segmented speech, simplifying the task and overlooking significant challenges. This narrow focus, coupled with widespread terminological inconsistencies, is limiting the applicability of research outcomes to real-world applications, ultimately hindering progress in the field. Our extensive literature review of 110 papers not only reveals these critical issues in current research but also serves as the foundation for our key contributions. We 1) define the steps and core components of a SimulST system, proposing a standardized terminology and taxonomy; 2) conduct a thorough analysis of community trends, and 3) offer concrete recommendations and future directions to bridge the gaps in existing literature, from evaluation frameworks to system architectures, for advancing the field towards more realistic and effective SimulST solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。