arXiv:2512.05100cs.CLcs.AI2025-12

让机器翻译文档时保持原有结构,效果显著提升。

Structured Document Translation via Format Reinforcement Learning

  • 用强化学习直接优化结构相似性和节点级翻译质量
  • 在SAP文档数据集上六项指标均提升,结构错误减少
  • 适合需要精准保留文档格式的工业级翻译场景

现有结构化文本翻译多局限于句子级别,难以处理复杂的文档级XML或HTML结构。为此,我们提出格式强化学习(FormatRL),在监督微调模型基础上,采用组相对策略优化方法,直接优化两类新型结构感知奖励:1)TreeSim,衡量预测与参考XML树之间的结构相似性;2)Node-chrF,评估XML节点级别的翻译质量。此外,引入细粒度指标StrucAUC,区分轻微错误与重大结构失败。在SAP软件文档基准上的实验表明,该方法在六项指标上均有提升,分析还揭示了不同奖励函数对结构与翻译质量改进的贡献。

原文摘要 · Abstract (English)

Recent works on structured text translation remain limited to the sentence level, as they struggle to effectively handle the complex document-level XML or HTML structures. To address this, we propose \textbf{Format Reinforcement Learning (FormatRL)}, which employs Group Relative Policy Optimization on top of a supervised fine-tuning model to directly optimize novel structure-aware rewards: 1) TreeSim, which measures structural similarity between predicted and reference XML trees and 2) Node-chrF, which measures translation quality at the level of XML nodes. Additionally, we apply StrucAUC, a fine-grained metric distinguishing between minor errors and major structural failures. Experiments on the SAP software-documentation benchmark demonstrate improvements across six metrics and an analysis further shows how different reward functions contribute to improvements in both structural and translation quality.

文档翻译强化学习结构生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。