哈工大深圳分校用Whisper+Krutrim实现印地语语音翻译,性能显著提升
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
- 端到端系统融合Whisper ASR与专用于印地语的LLM Krutrim
- 英→印地语方向平均BLEU达28.88,印地语→英语为27.86
- 链式思维可提升翻译质量,但输出格式稳定性不足
本文介绍哈工大深圳分校在IWSLT 2025印地语赛道的语音到文本翻译(ST)系统,针对英语与印地语之间的双向翻译任务。为提升低资源场景下的翻译质量,提出一种端到端系统,将预训练的Whisper自动语音识别(ASR)模型与专用于印地语的大型语言模型(LLM)Krutrim相结合。实验结果表明,该系统在英语→印地语方向平均达到28.88的BLEU得分,在印地语→英语方向达到27.86。此外,我们探索了链式思维(Chain-of-Thought, CoT)方法:尽管其在成功解析输出时表现出显著提升(如泰米尔语→英语方向提升13.84 BLEU),但模型难以稳定遵循所需的CoT输出格式。
原文摘要 · Abstract (English)
This paper presents HITSZ's submission for the IWSLT 2025 Indic track, focusing on speech-to-text translation (ST) for English-to-Indic and Indic-to-English language pairs. To enhance translation quality in this low-resource scenario, we propose an end-to-end system integrating the pre-trained Whisper automated speech recognition (ASR) model with Krutrim, an Indic-specialized large language model (LLM). Experimental results demonstrate that our end-to-end system achieved average BLEU scores of $28.88$ for English-to-Indic directions and $27.86$ for Indic-to-English directions. Furthermore, we investigated the Chain-of-Thought (CoT) method. While this method showed potential for significant translation quality improvements on successfully parsed outputs (e.g. a $13.84$ BLEU increase for Tamil-to-English), we observed challenges in ensuring the model consistently adheres to the required CoT output format.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。