提出新指标与优化框架,让文档翻译更忠实于原文结构。
STAR : Sentence Translation Alignment Rate for Document-to-Document Machine Translation

- 用句子对齐率量化翻译结构一致性
- 在新闻和文学领域显著提升翻译质量与结构完整性
- 小模型可超越大模型,适合资源受限场景
大型语言模型使机器翻译从句级转向文档到文档(Doc2Doc),有望提升全局连贯性。然而,单次生成常出现结构错位,表现为句子遗漏或幻觉,违背源文与目标文的对应要求。为此,我们提出句子翻译对齐率(STAR),一种显式衡量句级结构保真度的辅助指标。基于此,提出STAR掩码偏好优化(StarPO)框架,通过结构质量排序文档级假设,并使用动态对齐掩码聚焦优化错位段落。跨新闻与文学领域的实验表明,StarPO显著提升翻译质量与结构完整性。尤为突出的是,该方法使紧凑模型超越如GPT-4o等大型专有系统,同时保持更高令牌效率。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have enabled a shift from sentence-level to document-to-document (Doc2Doc) machine translation, promising improved global coherence. However, document-to-document generation in a single pass frequently suffers from structural misalignment, manifesting as sentence omissions or hallucinations that violate the core requirement of source-target correspondence. To address this, we introduce Sentence Translation Alignment Rate (STAR), an auxiliary metric that explicitly quantifies sentence-level structural fidelity. Building on this, we propose STAR-masked Preference Optimization (StarPO), a framework that ranks document-level hypotheses by structural quality and utilizes a dynamic alignment mask to focus optimization on misaligned segments. Experimental results across news and literary domains demonstrate that StarPO significantly enhances translation quality and structural integrity. Notably, StarPO allows compact models to surpass the performance of massive proprietary systems like GPT-4o while maintaining superior token efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。