arXiv:2505.21937cs.CL2025-05ACL被引 1

用动态图网络提升印地语等语言的习语翻译准确率

Graph-Assisted Culturally Adaptable Idiomatic Translation for Indic Languages

  • 构建自适应图神经网络,学习习语间的复杂映射关系
  • 在无资源条件下仍显著提升英译印地语习语质量
  • 适合需要跨文化精准翻译的低资源语言项目

习语和多词表达的翻译需深入理解源语言与目标语言的文化内涵。这一挑战因习语翻译的‘一对多’特性而加剧——同一源语言习语在不同文化背景和语境下可能对应多个目标语言表达。传统静态知识图谱和基于提示的方法难以捕捉此类复杂关系,常导致翻译效果不佳。为此,我们提出IdiomCE,一种基于自适应图神经网络(GNN)的方法,能够学习习语间复杂的映射关系,在训练过程中有效泛化至已见与未见节点。该方法在资源受限场景下仍能提升翻译质量,适用于小型模型。我们在多个习语翻译数据集上使用无参考指标评估,证明了其在英译多种印度语言时的显著改进。

原文摘要 · Abstract (English)

Translating multi-word expressions (MWEs) and idioms requires a deep understanding of the cultural nuances of both the source and target languages. This challenge is further amplified by the one-to-many nature of idiomatic translations, where a single source idiom can have multiple target-language equivalents depending on cultural references and contextual variations. Traditional static knowledge graphs (KGs) and prompt-based approaches struggle to capture these complex relationships, often leading to suboptimal translations. To address this, we propose IdiomCE, an adaptive graph neural network (GNN) based methodology that learns intricate mappings between idiomatic expressions, effectively generalizing to both seen and unseen nodes during training. Our proposed method enhances translation quality even in resource-constrained settings, facilitating improved idiomatic translation in smaller models. We evaluate our approach on multiple idiomatic translation datasets using reference-less metrics, demonstrating significant improvements in translating idioms from English to various Indian languages.

习语翻译图神经网络低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。