arXiv:2504.20451cs.CLcs.IR2025-04ACL被引 1

提升英韩跨文化翻译质量,发现大模型在实体翻译上仍存短板。

Team ACK at SemEval-2025 Task 2: Beyond Word-for-Word Machine Translation for English-Korean Pairs

  • 对比13种模型,发现大模型优于传统机器翻译系统。
  • 实体翻译错误率高,尤其涉及文化适配时表现不佳。
  • 构建错误分类体系,助力未来实现更自然的跨文化翻译。

在英韩双语间翻译知识密集型、实体丰富的文本时,需超越字面直译,保留语言与文化特异性。我们评估了13种模型(大语言模型与机器翻译系统),结合自动指标与双语标注员的人工评价。结果表明,大语言模型整体优于传统机器翻译系统,但在需要文化适配的实体翻译上表现不足。通过构建错误分类体系,识别出错误回应与实体名称错误为关键问题,且性能随实体类型与流行度变化。本研究揭示了现有自动评估指标的局限性,旨在推动更具文化敏感性的机器翻译发展。

原文摘要 · Abstract (English)

Translating knowledge-intensive and entity-rich text between English and Korean requires transcreation to preserve language-specific and cultural nuances beyond literal, phonetic or word-for-word conversion. We evaluate 13 models (LLMs and MT models) using automatic metrics and human assessment by bilingual annotators. Our findings show LLMs outperform traditional MT systems but struggle with entity translation requiring cultural adaptation. By constructing an error taxonomy, we identify incorrect responses and entity name errors as key issues, with performance varying by entity type and popularity level. This work exposes gaps in automatic evaluation metrics and hope to enable future work in completing culturally-nuanced machine translation.

机器翻译跨文化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。