arXiv:2606.05924cs.CLcs.AI2026-06ACL

用多维度生成数据提升文学翻译质量,效果优于真实参考译文。

Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach

论文配图:Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach
图 1 · 摘自论文原文
  • 通过专用大模型分维度生成翻译参考和偏好数据
  • 生成数据使SFT性能提升8.65点,最高达69.07分
  • 适合追求高质量文学翻译的开发者与研究者

文学翻译因高质量标注数据稀缺且需兼顾表达流畅性与文学效果而面临挑战。本文提出一种多维度迭代优化框架,利用针对不同质量维度的专用大模型生成高质量翻译参考与偏好数据,并用于监督微调与强化学习。实验表明,生成的参考译文在监督微调中比原始真值高出8.65 CEA100分;在强化学习中,采用显式奖励模型的GRPO比DPO表现更优,额外提升1.51分。这归因于两阶段训练的稳定性与GRPO的在线探索能力。所构建的LitMT-8B与LitMT-14B模型在MetaphorTrans英译中文学术基准上分别取得67.25与69.07分,性能媲美Claude Sonnet 4.5(68.43分),并展现出对域外文学作品(如欧·亨利)的强大泛化能力。

原文摘要 · Abstract (English)

Literary translation poses unique challenges due to the scarcity of high-quality annotated data and the need to balance expression fluency with literary effect. We present a multi-aspect iterative refinement framework that generates high-quality translation references and preference data through specialized LLM translators, each targeting a distinct quality dimension. We leverage the generated data for supervised fine-tuning and reinforcement learning. Experiments show that our generated references outperform the original ground truth for SFT by 8.65 CEA100 points. For reinforcement learning, we find that DPO leads to performance degradation in this setting, while leveraging an explicit reward model for GRPO yields an additional 1.51 point improvement. We attribute this to the stability of two-stage training and GRPO's online exploration capability. Our resulting models, LitMT-8B and LitMT-14B, achieve 67.25 and 69.07 CEA100 respectively on the MetaphorTrans English-to-Chinese literary translation benchmark, competitive with Claude Sonnet 4.5 at 68.43, and demonstrate strong generalization to out-of-domain literary work (i.e., O. Henry).

文学翻译大模型训练数据生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。