arXiv:2509.16375cs.CL2025-09EMNLP

用轻量适配器统一语音与文本翻译,提升多模态翻译效果。

Whisper-UT: A Unified Translation Framework for Speech and Text

  • 通过轻量适配器实现语音与文本任务的统一建模。
  • 在无三向平行数据条件下,性能优于基线模型。
  • 适合需要多模态翻译的场景,如语音转译、跨模态理解。

编码器-解码器模型在语音和文本任务中取得显著进展,但高效适应多种单/多模态场景仍是开放挑战。本文提出Whisper-UT,一种统一且高效的框架,利用轻量适配器实现任务间的无缝迁移,涵盖需同时依赖语音和源语言文本输入的多模态机器翻译(MMT)任务。通过将自动语音识别(ASR)结果或真实转录作为提示,该方法不仅支持双模态并行处理,还通过两阶段解码策略提升语音翻译(ST)性能。我们以Whisper模型为基础验证方法,但原则上可推广至其他多任务模型。实验表明,跨模态与跨任务微调有效提升性能,且无需三向平行数据。结果证明该框架在多模态翻译中具有高灵活性、高效性与通用性。

原文摘要 · Abstract (English)

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challenge. In this paper, we propose Whisper-UT, a unified and efficient framework that leverages lightweight adapters to enable seamless adaptation across tasks, including a multi-modal machine translation (MMT) task that explicitly conditions translation on both speech and source language text inputs. By incorporating ASR hypotheses or ground-truth transcripts as prompts, this approach not only enables the system to process both modalities simultaneously but also enhances speech translation (ST) performance through a 2-stage decoding strategy. We demonstrate our methods using the Whisper model, though in principle they are general and could be applied to similar multitask models. We highlight the effectiveness of cross-modal and cross-task fine-tuning, which improves performance without requiring 3-way parallel data. Our results underscore the flexibility, efficiency, and general applicability of the proposed framework for multi-modal translation.

多模态翻译语音翻译适配器Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。