首个专为网络小说翻译设计的多智能体评估框架,解决传统指标无法衡量叙事与文化忠实度的问题。
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
- 构建多智能体评估系统模拟专家讨论,超越词重合度衡量翻译质量
- 基于1.8万对专家标注语料,覆盖六维评估维度,发现中文训练模型表现更优
- 开源评估数据集与基准测试,适合研究网络小说翻译与LLM评估的学者
大型语言模型(LLMs)显著推进了机器翻译发展,但其在网文翻译中的实际效果仍不清晰。现有基准依赖表面指标,难以捕捉该类文本的独特属性。为此,我们提出DITING——首个针对网络小说翻译的综合性评估框架,从成语翻译、词汇歧义、术语本地化、时态一致性、零代词消解和文化安全六个维度进行评估,依托超过1.8万对专家标注的中英句对。我们进一步提出AgentEval,一种基于推理的多智能体评估框架,模拟专家协商过程,实现与人工判断最高相关性的自动评估,优于七种测试指标。为支持指标对比,我们构建MetricAlign,一个包含300对句子的元评估数据集,附有错误标签与量化评分。对十四种开放、闭源及商用模型的全面评估显示,中文训练的LLMs表现优于更大规模的外国模型,DeepSeek-V3在忠实度与风格一致性上最优。本工作确立了基于LLM的网文翻译新范式,并提供公开资源以推动后续研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on surface-level metrics that fail to capture the distinctive traits of this genre. To address these gaps, we introduce DITING, the first comprehensive evaluation framework for web novel translation, assessing narrative and cultural fidelity across six dimensions: idiom translation, lexical ambiguity, terminology localization, tense consistency, zero-pronoun resolution, and cultural safety, supported by over 18K expert-annotated Chinese-English sentence pairs. We further propose AgentEval, a reasoning-driven multi-agent evaluation framework that simulates expert deliberation to assess translation quality beyond lexical overlap, achieving the highest correlation with human judgments among seven tested automatic metrics. To enable metric comparison, we develop MetricAlign, a meta-evaluation dataset of 300 sentence pairs annotated with error labels and scalar quality scores. Comprehensive evaluation of fourteen open, closed, and commercial models reveals that Chinese-trained LLMs surpass larger foreign counterparts, and that DeepSeek-V3 delivers the most faithful and stylistically coherent translations. Our work establishes a new paradigm for exploring LLM-based web novel translation and provides public resources to advance future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。