arXiv:2508.19481cs.CLcs.AI2025-08AAAI被引 1

用词典+强化学习提升濒危语言翻译,西班牙语转瓦尤纳伊基语效果显著。

Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study

  • 让模型在生成时可选择查阅双语词典,结合监督微调与强化学习
  • 在美洲语料共享任务测试集上比无词典基线提升18%相对性能
  • 适合研究低资源语言、工具增强型AI或濒危语言保护的学者

低资源机器翻译仍是大语言模型(LLMs)的重大挑战,因这些语言在预训练中曝光不足且微调时并行数据有限。本文提出一种新方法,通过整合外部词典工具,并采用强化学习进行端到端训练,以提升低资源语言翻译质量。聚焦西班牙语-瓦尤纳伊基语这对语言,将翻译视为一个工具增强的决策问题,使模型可在生成过程中选择性查阅双语词典。该方法结合监督指令微调与引导奖励策略优化(GRPO),使模型学会何时以及如何有效使用工具。使用BLEU分数作为奖励信号引导学习过程。初步结果显示,工具增强模型在美洲语料共享任务2025的西班牙语-瓦尤纳伊基语测试集上,相比先前工作最高提升3.37 BLEU,相对于无词典访问的监督基线实现18%的相对增益。我们还通过消融实验评估了模型架构和训练策略的影响,对比Qwen2.5-0.5B-Instruct、LLaMA及此前基于NLLB的系统。结果表明,将大语言模型与外部工具结合,并利用强化学习,对低资源语言场景下的翻译质量提升具有巨大潜力。

原文摘要 · Abstract (English)

Low-resource machine translation remains a significant challenge for large language models (LLMs), which often lack exposure to these languages during pretraining and have limited parallel data for fine-tuning. We propose a novel approach that enhances translation for low-resource languages by integrating an external dictionary tool and training models end-to-end using reinforcement learning, in addition to supervised fine-tuning. Focusing on the Spanish-Wayuunaiki language pair, we frame translation as a tool-augmented decision-making problem in which the model can selectively consult a bilingual dictionary during generation. Our method combines supervised instruction tuning with Guided Reward Policy Optimization (GRPO), enabling the model to learn both when and how to use the tool effectively. BLEU similarity scores are used as rewards to guide this learning process. Preliminary results show that our tool-augmented models achieve up to +3.37 BLEU improvement over previous work, and a 18% relative gain compared to a supervised baseline without dictionary access, on the Spanish-Wayuunaiki test set from the AmericasNLP 2025 Shared Task. We also conduct ablation studies to assess the effects of model architecture and training strategy, comparing Qwen2.5-0.5B-Instruct with other models such as LLaMA and a prior NLLB-based system. These findings highlight the promise of combining LLMs with external tools and the role of reinforcement learning in improving translation quality in low-resource language settings.

低资源翻译强化学习词典增强濒危语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。