arXiv:2505.13036cs.CLcs.AI2025-05被引 5

KIT用大模型提升离线语音翻译与指令跟随性能

KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025

  • 多语音识别系统输出融合后,用大模型进行文档级上下文处理
  • 离线语音翻译任务达新高,指令跟随实现端到端多任务处理
  • 适合关注大模型在语音任务中应用的研究者和开发者

国际口语翻译研讨会(IWSLT)的范围已从传统语音翻译扩展至语音问答、摘要等任务,这得益于大型语言模型(LLMs)的发展。本文介绍了卡尔斯鲁厄理工学院在离线语音翻译(Offline ST)与指令跟随(IF)赛道的提交方案。针对离线ST任务,提出采用多个自动语音识别系统,其输出通过具备文档级上下文的大模型进行融合,并经两步翻译及额外优化步骤提升译文质量。对于IF任务,构建了融合语音编码器与大模型的端到端模型,支持多种指令任务,并通过文档级精修进一步提升输出质量。

原文摘要 · Abstract (English)

The scope of the International Workshop on Spoken Language Translation (IWSLT) has recently broadened beyond traditional Speech Translation (ST) to encompass a wider array of tasks, including Speech Question Answering and Summarization. This shift is partly driven by the growing capabilities of modern systems, particularly with the success of Large Language Models (LLMs). In this paper, we present the Karlsruhe Institute of Technology's submissions for the Offline ST and Instruction Following (IF) tracks, where we leverage LLMs to enhance performance across all tasks. For the Offline ST track, we propose a pipeline that employs multiple automatic speech recognition systems, whose outputs are fused using an LLM with document-level context. This is followed by a two-step translation process, incorporating additional refinement step to improve translation quality. For the IF track, we develop an end-to-end model that integrates a speech encoder with an LLM to perform a wide range of instruction-following tasks. We complement it with a final document-level refinement stage to further enhance output quality by using contextual information.

语音翻译大模型指令跟随

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。