arXiv:2606.12911cs.CL2026-06中稿 · INTERSPEECH 2026

针对越南语语音翻译中语音识别错误导致的性能下降,提出基于发音相似性的数据增强方法。

PiDA: Phonetically-Informed Data Augmentation for Robust Vietnamese Speech Translation

  • 通过发音相似性嵌入生成类语音识别错误的噪声数据
  • 在错误语音识别输出上提升最高2.04的BLEU分数
  • 适合需要提升鲁棒性的语音翻译系统开发者

级联式语音翻译系统在自动语音识别(ASR)输出错误文本时会面临误差传播问题。本文首次对越南语语音翻译中的ASR错误进行系统分类,按发音原因对替换错误进行归类,并使用线性混合效应模型量化其对下游神经机器翻译(NMT)性能的影响。研究确认,多数ASR替换错误源于发音混淆而非随机噪声,且这些发音错误显著降低语音翻译质量。基于此发现,提出发音感知数据增强(PiDA),通过发音词嵌入将词语替换为发音相似的替代词,生成类ASR噪声。在FLEURS越南语-英语数据集上对PiDA增强版本进行微调,可使错误语音识别输出的翻译性能提升最多2.04 BLEU,同时轻微改善干净文本表现。

原文摘要 · Abstract (English)

Cascaded speech translation (ST) systems suffer from error propagation when Automatic Speech Recognition (ASR) outputs incorrect transcripts. We present the first systematic categorization of ASR errors for Vietnamese ST, classifying substitution errors by phonetic cause and quantifying their impact on downstream Neural Machine Translation (NMT) performance using Linear Mixed-Effects Modelling. We confirm that most ASR substitution errors arise from phonetic confusions rather than random noise, and that these phonetic errors significantly degrade ST quality. Motivated by this finding, we propose Phonetically-Informed Data Augmentation (PiDA), which generates ASR-like corruptions by substituting words with phonetically similar alternatives using phonetic word embeddings. Fine-tuning on a PiDA-augmented version of FLEURS Vietnamese-English improves translation of erroneous ASR outputs (up to +2.04 BLEU over standard fine-tuning) while also slightly improving clean-text performance.

语音翻译数据增强越南语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。