用音素增强思维链,让低资源语言也能实现语音翻译。
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
- 将音素识别作为中间步骤,提升跨语言迁移能力
- 在低资源场景下显著提升翻译质量,支持零样本翻译
- 适合研究多语言语音翻译与资源匮乏语言的学者
我们提出一种语音到文本翻译(S2TT)方法,将音素表示引入思维链(CoT)框架,以提升低资源和零资源场景下的翻译效果。通过引入音素识别作为中间步骤,增强了跨语言迁移能力,使无标注语音数据的语言也能实现翻译。系统基于多语言大模型构建,并扩展其处理语音与音素的能力。训练采用渐进式课程学习策略,逐步增加任务复杂度。在多语言S2TT基准上的实验表明,音素增强的CoT在低资源条件下提升了翻译质量,实现了零资源翻译,仅轻微影响高资源性能。尽管存在这一权衡,结果表明音素驱动的CoT是迈向更广泛语言覆盖S2TT的重要一步。
原文摘要 · Abstract (English)
We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing phoneme recognition as an intermediate step, we enhance cross-lingual transfer, enabling translation even for languages with no labeled speech data. Our system builds on a multilingual LLM, which we extend to process speech and phonemes. Training follows a curriculum learning strategy that progressively introduces more complex tasks. Experiments on multilingual S2TT benchmarks show that phoneme-augmented CoT improves translation quality in low-resource conditions and enables zero-resource translation, while slightly impacting high-resource performance. Despite this trade-off, our findings demonstrate that phoneme-based CoT is a promising step toward making S2TT more accessible across diverse languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。