arXiv:2512.07265cs.CLeess.AS2025-12被引 1

构建首个泰卢固语-英语语音翻译基准,验证端到端模型在低资源下的竞争力

TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

  • 基于46小时人工校验数据构建泰卢固语-英语语音翻译基准
  • 端到端模型仅用少量泰卢固语数据即接近级联模型性能
  • 发现传统评估指标比BERTScore更适合泰卢固语翻译评测

尽管泰卢固语有超过8000万使用者,但针对该形态丰富的语言的语音翻译研究仍严重不足。本文通过46小时经人工验证的跨语言语音转录数据(训练/验证/测试集分别为30小时、8小时、8小时),构建了高质量的泰卢固语-英语语音翻译基准。系统对比级联式与端到端架构发现,虽然基于泰卢固语特训数据的IndicWhisper + IndicMT表现最优,但微调后的SeamlessM4T模型在使用远少于100小时的泰卢固语数据下也展现出显著竞争力。这表明,在充分调参和足够并行数据支持下,端到端系统可在低资源场景中达到与级联方法相当的性能。此外,对BLEU、METEOR、ChrF++、ROUGE-L、TER和BERTScore等指标的可靠性研究显示,传统指标在区分泰卢固语-英语翻译质量方面优于BERTScore。本工作贡献包括:可复现的泰卢固语-英语语音翻译基准、端到端系统在低资源下具备竞争潜力的实证证据,以及针对形态复杂语言对的自动评估实践指导。

原文摘要 · Abstract (English)

Despite Telugu being spoken by over 80 million people, speech translation research for this morphologically rich language remains severely underexplored. We address this gap by developing a high-quality Telugu--English speech translation benchmark from 46 hours of manually verified CSTD corpus data (30h/8h/8h train/dev/test split). Our systematic comparison of cascaded versus end-to-end architectures shows that while IndicWhisper + IndicMT achieves the highest performance due to extensive Telugu-specific training data, finetuned SeamlessM4T models demonstrate remarkable competitiveness despite using significantly less Telugu-specific training data. This finding suggests that with careful hyperparameter tuning and sufficient parallel data (potentially less than 100 hours), end-to-end systems can achieve performance comparable to cascaded approaches in low-resource settings. Our metric reliability study evaluating BLEU, METEOR, ChrF++, ROUGE-L, TER, and BERTScore against human judgments reveals that traditional metrics provide better quality discrimination than BERTScore for Telugu--English translation. The work delivers three key contributions: a reproducible Telugu--English benchmark, empirical evidence of competitive end-to-end performance potential in low-resource scenarios, and practical guidance for automatic evaluation in morphologically complex language pairs.

语音翻译低资源多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。