arXiv:2409.15911cs.CLcs.SD2024-09中稿 · ICASSP 2025被引 2

提出模块化梯度冲突缓解策略,提升实时语音翻译效率与内存利用率。

A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation

  • 在模块级别检测并用梯度投影解决多任务学习中的优化冲突
  • 在中高延迟下显著提升性能,离线任务提升0.68 BLEU
  • 相比其他方法降低超95%显存占用,适合资源受限场景

同时性语音翻译(SimulST)需在持续接收流式语音输入时生成目标语言文本,面临严峻的实时性挑战。尽管多任务学习常用于提升性能,但主任务与辅助任务间常出现优化冲突,影响整体效率。现有模型级冲突缓解方法不适用于此任务,加剧了计算低效并导致高显存消耗。为此,我们提出模块化梯度冲突缓解(MGCM)策略,在更细粒度的模块层级检测冲突,并通过梯度投影进行解决。实验表明,MGCM显著提升SimulST性能,尤其在中高延迟条件下,离线任务实现0.68 BLEU得分提升。同时,相较于其他冲突缓解方法,显存消耗降低超过95%,展现出对SimulST任务的强大适应性。

原文摘要 · Abstract (English)

Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95\% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.

语音翻译多任务学习梯度优化低显存

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。