突破翻译评估局限,实现长篇文档端到端自动评测
Extending Automatic Machine Translation Evaluation to Book-Length Documents
- 将文档视为连续文本,动态分割与对齐句子
- 在长文档上表现显著优于现有方法,接近人工标注精度
- 揭示多数开源大模型在宣称上下文长度下无法有效翻译
尽管大型语言模型(LLMs)在翻译性能和长上下文处理方面表现优异,但评估方法仍受限于句级评估,原因包括数据集限制、度量指标的令牌数约束以及严格的句边界要求。我们提出SEGALE评估方案,通过将文档视为连续文本,并应用句子分割与对齐方法,扩展现有自动评估指标至长文档翻译场景。该方法可处理任意长度的文档级提示生成翻译,同时识别漏译、多译及不规则句界问题。实验表明,该方案显著优于现有长文档评估方法,且与基于真实句对齐的评估结果相当。此外,我们将其应用于书本级文本,首次揭示许多开源大模型在声称的最大上下文长度下无法有效完成翻译任务。
原文摘要 · Abstract (English)
Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentence-level assessment due to dataset limitations, token number restrictions in metrics, and rigid sentence boundary requirements. We introduce SEGALE, an evaluation scheme that extends existing automatic metrics to long-document translation by treating documents as continuous text and applying sentence segmentation and alignment methods. Our approach enables previously unattainable document-level evaluation, handling translations of arbitrary length generated with document-level prompts while accounting for under-/over-translations and varied sentence boundaries. Experiments show our scheme significantly outperforms existing long-form document evaluation schemes, while being comparable to evaluations performed with groundtruth sentence alignments. Additionally, we apply our scheme to book-length texts and newly demonstrate that many open-weight LLMs fail to effectively translate documents at their reported maximum context lengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。