构建中文-英文专业领域篇章级翻译基准,揭示大模型与人类专家的显著差距。
DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains
- 设计跨7个专业领域的200篇长文本(平均1700词)篇章级翻译数据集
- 提出无参考的Metric-S评估系统,对准确、流畅、得体性评估与人工高度一致
- 实测顶尖大模型仍远低于人类专家,凸显专业翻译任务的挑战
专业领域中的篇章级翻译评估仍不充分,尽管其在知识传播与跨语言学术交流中至关重要。此类翻译需兼顾篇章连贯性与术语精确性,但现有评估方法多聚焦片段级的准确性和流畅性。为此,我们提出DiscoX,一个面向中文-英文专业领域篇章级翻译的新基准,包含7个领域共200篇由专业人士精心挑选的文本,平均长度超过1700个标记。为评估在DiscoX上的表现,我们还开发了Metric-S——一种无参考的自动评估系统,可细粒度衡量准确性、流畅性和得体性。Metric-S与人工判断具有强一致性,显著优于现有指标。实验发现,即使最先进的大模型在该任务上仍明显落后于人类专家。这一结果验证了DiscoX的难度,凸显了实现专业级机器翻译的持续挑战。所提出的基准与评估体系为更严谨的评测提供了坚实框架,助力未来基于大模型的翻译发展。
原文摘要 · Abstract (English)
The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. While these translations demand discourse-level coherence and strict terminological precision, current evaluation methods predominantly focus on segment-level accuracy and fluency. To address this limitation, we introduce DiscoX, a new benchmark for discourse-level and expert-level Chinese-English translation. It comprises 200 professionally-curated texts from 7 domains, with an average length exceeding 1700 tokens. To evaluate performance on DiscoX, we also develop Metric-S, a reference-free system that provides fine-grained automatic assessments across accuracy, fluency, and appropriateness. Metric-S demonstrates strong consistency with human judgments, significantly outperforming existing metrics. Our experiments reveal a remarkable performance gap: even the most advanced LLMs still trail human experts on these tasks. This finding validates the difficulty of DiscoX and underscores the challenges that remain in achieving professional-grade machine translation. The proposed benchmark and evaluation system provide a robust framework for more rigorous evaluation, facilitating future advancements in LLM-based translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。