arXiv:2601.03790cs.CLcs.AI2026-01ACL被引 3

针对新造词的机器翻译,提出基于检索增强的智能代理框架。

NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning

  • 构建16语言、75方向的新造词翻译数据集与维基词典检索工具
  • 利用强化学习训练翻译智能体,提升新造词翻译准确率
  • 设计新奖励机制与自适应生成策略,适合语言研究者与翻译系统开发者

新造词感知的机器翻译旨在将包含新造词的源句翻译为目标语言,该领域相较于通用机器翻译仍较薄弱。本文提出一种基于维基词典检索工具的智能体框架NeoAMT。首先,基于约1000万条英文维基词典数据构建涵盖16种语言、75个翻译方向的专用数据集,并从相同数据中清洗出约300万条记录作为检索语料库。随后,利用该数据集和检索工具,通过强化学习训练翻译智能体并评估其准确性。进一步提出一种新型奖励设计与自适应滚动生成策略,结合翻译难度动态优化生成过程,显著提升翻译质量。

原文摘要 · Abstract (English)

Neologism-aware machine translation aims to translate source sentences containing neologisms into target languages. This field remains underexplored compared with general machine translation (MT). In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit. Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary. The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump. The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump. We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologism-aware machine translation. Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits translation difficulty to further improve the translation quality of translation agents using our search toolkit.

机器翻译新造词强化学习智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。