arXiv:2508.11189cs.CLcs.SD2025-08

提出寄生双尺度模型,提升多语言语音翻译效率与精度

Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation

  • 用寄生式双尺度架构结合推测采样与知识蒸馏
  • 在6种语言上达最优性能,推理速度提升2.6倍
  • 适合需要高效本地部署的多语言语音翻译场景

近年来,语音转文本翻译取得了进展,涌现出可同时处理多语言对的统一模型。然而,这些模型通常参数量大,难以在本地部署中兼顾推理效率与性能。本文提出一种创新的寄生双尺度方法,结合增强的推测采样、模型压缩与知识蒸馏技术。基于Whisper Medium模型,我们构建了whisperM2M模型,并引入新颖的KVSPN模块,在六种主流语言上实现当前最佳(SOTA)性能,同时提升推理效率。KVSPN实现40%的速度提升且不损失BLEU得分;结合蒸馏后,相较原始Whisper Medium模型提速2.6倍,性能更优。

原文摘要 · Abstract (English)

Recent advancements in speech-to-text translation have led to the development of multilingual models capable of handling multiple language pairs simultaneously. However, these unified models often suffer from large parameter sizes, making it challenging to balance inference efficiency and performance, particularly in local deployment scenarios. We propose an innovative Parasitic Dual-Scale Approach, which combines an enhanced speculative sampling method with model compression and knowledge distillation techniques. Building on the Whisper Medium model, we enhance it for multilingual speech translation into whisperM2M, and integrate our novel KVSPN module, achieving state-of-the-art (SOTA) performance across six popular languages with improved inference efficiency. KVSPN enables a 40\% speedup with no BLEU score degradation. Combined with distillation methods, it represents a 2.6$\times$ speedup over the original Whisper Medium with superior performance.

语音翻译模型压缩多语言推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。