用图结构增强大模型,让文档翻译更连贯准确。
GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation
- 构建有向无环图建模文档内语义依赖关系
- 在8个翻译方向上平均提升2.8 dBLEU
- 适合需要高一致性翻译的场景
文档级机器翻译常难以有效捕捉语篇层面现象。现有方法多依赖启发式规则分割文档为语篇单元,往往与真实语篇结构不符,且翻译过程中难以保持一致性。为此,我们提出图增强型代理框架GRAFT,利用大语言模型代理实现文档翻译。该框架融合分段、基于有向无环图的依赖建模与语篇感知翻译,形成统一系统。在八个翻译方向和六个不同领域的实验中,GRAFT显著优于当前最优文档级翻译系统。具体而言,在IWSLT2017的TED测试集上,相比强基线平均提升2.8 dBLEU;英译中领域特定任务中提升2.3 dBLEU。分析表明,GRAFT具备稳定处理语篇现象的能力,生成连贯且上下文准确的翻译结果。
原文摘要 · Abstract (English)
Document level Machine Translation (DocMT) approaches often struggle with effectively capturing discourse level phenomena. Existing approaches rely on heuristic rules to segment documents into discourse units, which rarely align with the true discourse structure required for accurate translation. Otherwise, they fail to maintain consistency throughout the document during translation. To address these challenges, we propose Graph Augmented Agentic Framework for Document Level Translation (GRAFT), a novel graph based DocMT system that leverages Large Language Model (LLM) agents for document translation. Our approach integrates segmentation, directed acyclic graph (DAG) based dependency modelling, and discourse aware translation into a cohesive framework. Experiments conducted across eight translation directions and six diverse domains demonstrate that GRAFT achieves significant performance gains over state of the art DocMT systems. Specifically, GRAFT delivers an average improvement of 2.8 d BLEU on the TED test sets from IWSLT2017 over strong baselines and 2.3 d BLEU for domain specific translation from English to Chinese. Moreover, our analyses highlight the consistent ability of GRAFT to address discourse level phenomena, yielding coherent and contextually accurate translations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。