arXiv:2511.07230cs.CLcs.AI2025-11Conference of the …被引 2

用结构化语篇图提升长文档翻译质量,减少计算开销。

Discourse Graph Guided Document Translation with Large Language Models

  • 构建语篇图显式建模段落间关系,精准定位上下文。
  • 在六语言多领域测试中,翻译质量与术语一致性均超基线。
  • 适合需要高连贯性长文本翻译的场景,如法律、科技文档。

将大语言模型应用于全文翻译仍面临长距离依赖捕捉难、语篇连贯性难以保持的挑战。尽管近期基于代理的机器翻译系统通过多代理协同和持久记忆缓解了上下文窗口限制,但其需大量计算资源且对记忆检索策略敏感。我们提出TransGraph框架,通过结构化的语篇图显式建模段落间的关联,并仅对相关图邻域进行条件约束,而非依赖顺序或全量上下文。在涵盖六种语言、多个领域的三个文档级机器翻译基准上,TransGraph始终优于强基线,在翻译质量与术语一致性方面表现优异,同时显著降低令牌开销。

原文摘要 · Abstract (English)

Adapting large language models to full document translation remains challenging due to the difficulty of capturing long-range dependencies and preserving discourse coherence throughout extended texts. While recent agentic machine translation systems mitigate context window constraints through multi-agent orchestration and persistent memory, they require substantial computational resources and are sensitive to memory retrieval strategies. We introduce TransGraph, a discourse-guided framework that explicitly models inter-chunk relationships through structured discourse graphs and selectively conditions each translation segment on relevant graph neighbourhoods rather than relying on sequential or exhaustive context. Across three document-level MT benchmarks spanning six languages and diverse domains, TransGraph consistently surpasses strong baselines in translation quality and terminology consistency while incurring significantly lower token overhead.

文档翻译语篇图LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。