arXiv:2505.15333cs.CLcs.AI2025-05ACL被引 2

用单元语言引导语音建模,提升无文本语音翻译效果

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

  • 引入单元语言作为类文本表示,指导语音特征提取
  • 在Voxpupil上四语言实验,性能接近有文本模型
  • 提出任务提示机制缓解源目标语言冲突

文本无依赖的语音到语音翻译(S2ST)虽取得进展,但仍面临两大挑战:跨模态(CM)——从不同语音信号中提取语言特征;跨语言(CL)——长序列中不同语言间的对齐学习。本文提出单元语言概念,通过n-gram语言建模构建类文本表示形式,用于指导语音建模。采用多任务学习融合单元语言信息。初步实验发现同时使用源语言和目标语言单元时存在冲突,因此提出任务提示建模以缓解。在Voxpupil数据集的四种语言上进行实验,结果表明该方法显著优于强基线模型,性能接近使用文本训练的模型。

原文摘要 · Abstract (English)

The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.

语音翻译无文本语言建模多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。