用可验证奖励激励模型挖掘参数中的文化适配翻译知识。
Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation

- 设计基于实体的可验证奖励,引导模型学习文化适配的推理过程。
- 仅用7千样本训练,使大模型在5万未见实体上准确率提升8.21个百分点。
- 适用于需要跨文化精准翻译的场景,尤其适合资源有限时的微调。
跨文化实体翻译对大语言模型仍具挑战,常生成字面或音译结果而非文化适配翻译。但相关知识可能已编码于模型参数中。为激励有效利用这些参数知识,我们提出EA-RLVR(实体锚定强化学习与可验证奖励),一种无需外部知识库的训练框架。该框架以可验证的实体级奖励信号为监督,并引入轻量结构门控稳定优化过程。这一设计促使模型学习稳健的推理机制,而非简单模仿参考翻译。在XC-Translate数据集上评估显示,仅用7千样本训练,即可将Qwen3-14B在包含5万完全未见实体的测试集上的实体翻译准确率从23.66%提升至31.87%。所学翻译能力还迁移至通用翻译任务,在WMT24++上带来+1.35 XCOMET提升,扩展优化后达+1.59。对$pass@k$动态和奖励设计的分析表明,性能提升源于更优的采样效率和稳定的优化空间。
原文摘要 · Abstract (English)
Cross-cultural entity translation remains challenging for large language models (LLMs) as literal or phonetic renderings are usually yielded instead of culturally appropriate translations in context. However, relevant knowledge may already be encoded in model parameters during large-scale pre-training. To incentivize the effective use of parametric knowledge, we propose EA-RLVR (Entity-Anchored Reinforcement Learning with Verifiable Rewards), a training framework that optimizes cross-cultural entity translation without relying on external knowledge bases. EA-RLVR anchors supervision on a verifiable, entity-level reward signal and incorporates lightweight structural gates to stabilize optimization. This design steers the model toward learning a robust reasoning process rather than merely imitating reference translations. We evaluate EA-RLVR on XC-Translate and observe consistent improvements in both entity translation accuracy and out-of-domain generalization. Specifically, training on merely 7k samples boosts Qwen3-14B's entity translation accuracy from 23.66\% to 31.87\% on a 50k test set comprising entirely unseen entities. The learned entity translation ability also transfers to general translation, yielding +1.35 XCOMET on WMT24++, which scales to +1.59 with extended optimization. Extensive analyses of $pass@k$ dynamics and reward formulations attribute these gains to superior sampling efficiency and a stable optimization landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。