arXiv:2506.02995cs.CL2025-06ACL被引 2

语音转写系统翻译成语时表现差,需专门优化。

It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems

  • 对比文本与语音转写系统,评估成语翻译效果。
  • 语音转写系统成语翻译准确率显著下降,常译成字面意思。
  • 大模型和文本翻译系统更擅长处理成语,适合研究者参考。

成语是意义不可从字面推断的词组。尽管现代机器翻译系统取得显著进展,但成语翻译仍是重大挑战,尤其在语音转写系统中,相关研究尤为稀缺。本文系统评估了在德语-英语、俄语-英语双语对下,文本到文本机器翻译(MT)与语音到文本翻译(SLT)系统中成语翻译的表现。比较了先进的端到端SLT系统(SeamlessM4T SLT-to-text、Whisper Large v3)、MT系统(SeamlessM4T、No Language Left Behind)、大语言模型(DeepSeek、LLaMA)及级联方案。结果表明,SLT系统在成语数据上性能明显下降,即使在深层仍倾向于字面翻译;而MT系统与大语言模型对成语处理更优。这凸显了在SLT架构中引入成语专用策略与改进内部表示的必要性。

原文摘要 · Abstract (English)

Idioms are defined as a group of words with a figurative meaning not deducible from their individual components. Although modern machine translation systems have made remarkable progress, translating idioms remains a major challenge, especially for speech-to-text systems, where research on this topic is notably sparse. In this paper, we systematically evaluate idiom translation as compared to conventional news translation in both text-to-text machine translation (MT) and speech-to-text translation (SLT) systems across two language pairs (German to English, Russian to English). We compare state-of-the-art end-to-end SLT systems (SeamlessM4T SLT-to-text, Whisper Large v3) with MT systems (SeamlessM4T SLT-to-text, No Language Left Behind), Large Language Models (DeepSeek, LLaMA) and cascaded alternatives. Our results reveal that SLT systems experience a pronounced performance drop on idiomatic data, often reverting to literal translations even in higher layers, whereas MT systems and Large Language Models demonstrate better handling of idioms. These findings underscore the need for idiom-specific strategies and improved internal representations in SLT architectures.

成语翻译语音转写机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。