用语义约束提升低质量语音翻译的稳定性与准确性
Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels
- 引入语义正则化,通过冻结文本嵌入引导声学编码器
- 在30小时非专业标注数据上达到接近100%数据预训练模型效果
- 适合数据少、噪声大场景下的端到端语音翻译任务
端到端语音翻译在目标转录存在高方差和语义模糊时,常出现收敛慢、性能差的问题。本文提出听、注意、理解(LAU)语义正则化技术,在训练中约束声学编码器的隐空间。利用冻结的文本嵌入提供方向性辅助损失,使声学表示具备语言基础,且不增加推理开销。在30小时巴马拉语到法语的非专业标注数据集上评估,实验表明LAU模型在标准指标上表现可媲美使用100%更多数据预训练的E2E-ST系统,同时更有效地保留语义含义。此外,我们提出总参数漂移作为度量,证明语义约束会主动重构编码器权重,优先关注语义而非字面发音。研究显示LAU是后处理重评分的有效替代方案,尤其适用于数据稀缺或噪声较大的端到端语音翻译场景。
原文摘要 · Abstract (English)
End-to-End Speech Translation often shows slower convergence and worse performance when target transcriptions exhibit high variance and semantic ambiguity. We propose Listen, Attend, Understand (LAU), a semantic regularization technique that constrains the acoustic encoder's latent space during training. By leveraging frozen text embeddings to provide a directional auxiliary loss, LAU injects linguistic groundedness into the acoustic representation without increasing inference cost. We evaluate our method on a Bambara-to-French dataset with 30 hours of Bambara speech translated by non-professionals. Experimental results demonstrate that LAU models achieve comparable performance by standard metrics compared to an E2E-ST system pretrained with 100\% more data and while performing better in preserving semantic meaning. Furthermore, we introduce Total Parameter Drift as a metric to quantify the structural impact of regularization to demonstrate that semantic constraints actively reorganize the encoder's weights to prioritize meaning over literal phonetics. Our findings suggest that LAU is a robust alternative to post-hoc rescoring and a valuable addition to E2E-ST training, especially when training data is scarce and/or noisy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。