arXiv:2409.16539cs.AI2024-09被引 6

针对文学翻译中的上下文与风格保持难题,提出渐进式解码框架。

Context-aware and Style-related Incremental Decoding framework for Discourse-Level Literary Translation

论文配图:Context-aware and Style-related Incremental Decoding framework for Discourse-Level Literary Translation
图 1 · 摘自论文原文
  • 采用渐进解码机制,逐句翻译时考虑全文语境。
  • 在句子级和文档级上均提升BLEU分数,效果显著。
  • 适合需要保持文风一致性的文学翻译任务。

本文介绍了参加WMT24话语级文学翻译任务(中文-英文,受限赛道)的方法。文学翻译因语义细腻、习语丰富及叙事结构复杂而极具挑战。为此,我们基于中文-LLaMA2模型,结合持续预训练(CPT)与有监督微调(SFT)进行优化。提出一种新型渐进解码框架,使每句话的翻译都考虑上下文,保持文本连贯性与风格一致性,有效捕捉长距离依赖与文体特征,生成忠实还原原作文学质量的译文。实验表明,该方法在句子级与文档级均显著提升BLEU得分,验证了其在处理话语级文学翻译复杂性上的有效性。

原文摘要 · Abstract (English)

This report outlines our approach for the WMT24 Discourse-Level Literary Translation Task, focusing on the Chinese-English language pair in the Constrained Track. Translating literary texts poses significant challenges due to the nuanced meanings, idiomatic expressions, and intricate narrative structures inherent in such works. To address these challenges, we leveraged the Chinese-Llama2 model, specifically enhanced for this task through a combination of Continual Pre-training (CPT) and Supervised Fine-Tuning (SFT). Our methodology includes a novel Incremental Decoding framework, which ensures that each sentence is translated with consideration of its broader context, maintaining coherence and consistency throughout the text. This approach allows the model to capture long-range dependencies and stylistic elements, producing translations that faithfully preserve the original literary quality. Our experiments demonstrate significant improvements in both sentence-level and document-level BLEU scores, underscoring the effectiveness of our proposed framework in addressing the complexities of document-level literary translation.

文学翻译渐进解码上下文建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。