用语言类型学结构提升多语种语音翻译的效率与性能
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
- 将语言标签升级为类型学先验,构建分层语言表征
- 在多个指标上优于现有方法,3小时数据即达良好效果
- 适合数据稀缺场景下的多语种语音翻译研究
基于语音大语言模型的组合式语音到语音翻译(S2ST)系统近期表现优异。然而,现有系统或忽略源语言信息,或采用语言标签式的扁平嵌入表示,忽略了跨语言间共享的系统性语言结构,导致在标注数据稀缺时难以实现高效多语种适应。为此,我们提出S2ST-Omni 2,一种从扁平语言标签转向结构化类型学先验的多对一组合式S2ST框架。具体包括:基于类型学的分层语言编码以实现结构化源语言表征、动态门控的语言感知双CTC用于内容自适应声学调制,以及类型学感知的LLM提示用于解码端语言引导。在CVSS-C数据集上的实验表明,S2ST-Omni 2在BLEU、COMET、ASR-BLEU和BLASER 2.0等多个指标上均优于代表性S2ST方法。消融实验显示,三个层级策略具有互补优势。受控数据预算分析及仅使用约3小时监督数据的日语→英语评估进一步证明,显式类型学先验能为数据高效的多语种S2ST提供有效归纳偏置。
原文摘要 · Abstract (English)
Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST systems often either neglect source-language information or encode it through a language-as-label paradigm, representing each source language as an independent flat embedding. Such a design overlooks systematic linguistic structure shared across languages, which may limit data-efficient multilingual adaptation when supervised S2ST data are scarce. To address this issue, we propose S2ST-Omni 2, a many-to-one compositional S2ST framework that systematically reformulates multilingual language conditioning from flat language labels to structured typological priors. Specifically, S2ST-Omni 2 revisits language conditioning at three levels: typology-informed hierarchical language encoding for structured source-language representation, dynamically-gated language-aware Dual-CTC for content-adaptive acoustic modulation, and typology-aware LLM prompting for decoder-side linguistic guidance. Experiments on CVSS-C show that S2ST-Omni 2 achieves superior average performance among representative S2ST approaches across BLEU, COMET, ASR-BLEU, and BLASER 2.0 under the adopted evaluation protocol. Ablation studies indicate that the proposed representation-level, acoustic-level, and decoding-level strategies provide complementary benefits. Moreover, controlled data-budget analyses and a Japanese-to-English evaluation using only approximately 3 hours of supervised training data suggest that explicit typological priors provide useful inductive biases for data-efficient multilingual S2ST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。