用多智能体系统评估文学翻译质量,更懂风格与叙事
MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment
- 设计多智能体框架,分别从术语、叙事、风格三方面评估译文
- 在《小王子》等文本上,顶尖模型得分达0.890,优于传统指标
- 适合翻译研究者和需高保真译文的场景
文学翻译需保留文化细节与文体特征,但传统指标如BLEU和METEOR因侧重词汇重合,无法有效评估叙事连贯性与风格忠实度。为此,我们提出MAS-LitEval,一种基于大语言模型的多智能体系统,从术语、叙事和风格维度评估译文质量。我们在《小王子》和《亚瑟王宫廷中的康涅狄格州佬》的多种LLM译本上测试该系统,并与传统指标对比。结果表明,MAS-LitEval表现更优,顶尖模型在捕捉文学细节方面的得分高达0.890。本工作构建了一个可扩展、细致的翻译质量评估框架,为译者和研究人员提供实用工具。
原文摘要 · Abstract (English)
Literary translation requires preserving cultural nuances and stylistic elements, which traditional metrics like BLEU and METEOR fail to assess due to their focus on lexical overlap. This oversight neglects the narrative consistency and stylistic fidelity that are crucial for literary works. To address this, we propose MAS-LitEval, a multi-agent system using Large Language Models (LLMs) to evaluate translations based on terminology, narrative, and style. We tested MAS-LitEval on translations of The Little Prince and A Connecticut Yankee in King Arthur's Court, generated by various LLMs, and compared it to traditional metrics. \textbf{MAS-LitEval} outperformed these metrics, with top models scoring up to 0.890 in capturing literary nuances. This work introduces a scalable, nuanced framework for Translation Quality Assessment (TQA), offering a practical tool for translators and researchers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。