构建首个2024-2025年跨语言错误检测数据集,提升机器翻译安全评估精度。
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
- 基于2024-2025年真实来源构建1万对英德语句,含黄金与银色标注
- 平衡误差与非误差样本,支持细粒度错误分类与风险分析
- 适合关注翻译安全性、智能设备部署的NLP研究者使用
机器翻译中的关键错误检测(CED)旨在判断译文是否可用,是否存在不可接受的意义偏差。尽管WMT21英德语CED数据集提供了首个基准,但其规模小、标签不平衡、领域覆盖窄且时效性不足。本文提出SynCED-EnDe,包含1,000对黄金标注和8,000对银色标注的句子对,误差与非误差样本比例为50/50。数据源涵盖2024–2025年的真实语料(StackExchange、GOV.UK),引入显式错误子类、结构化触发标记及细粒度辅助判断(明显性、严重性、定位复杂度、上下文依赖性、适切性偏差)。这些设计支持超越二元检测的系统性误差风险与复杂度分析。数据集永久托管于GitHub与Hugging Face,附带文档、标注指南与基线脚本。在XLM-R等编码器上的基准实验显示,相比WMT21,性能显著提升,得益于标签平衡与精细化标注。我们期望SynCED-EnDe成为社区资源,推动机器翻译在信息检索与对话助手中的安全部署,尤其适用于可穿戴AI等新兴场景。
原文摘要 · Abstract (English)
Critical Error Detection (CED) in machine translation aims to determine whether a translation is safe to use or contains unacceptable deviations in meaning. While the WMT21 English-German CED dataset provided the first benchmark, it is limited in scale, label balance, domain coverage, and temporal freshness. We present SynCED-EnDe, a new resource consisting of 1,000 gold-labeled and 8,000 silver-labeled sentence pairs, balanced 50/50 between error and non-error cases. SynCED-EnDe draws from diverse 2024-2025 sources (StackExchange, GOV.UK) and introduces explicit error subclasses, structured trigger flags, and fine-grained auxiliary judgments (obviousness, severity, localization complexity, contextual dependency, adequacy deviation). These enrichments enable systematic analyses of error risk and intricacy beyond binary detection. The dataset is permanently hosted on GitHub and Hugging Face, accompanied by documentation, annotation guidelines, and baseline scripts. Benchmark experiments with XLM-R and related encoders show substantial performance gains over WMT21 due to balanced labels and refined annotations. We envision SynCED-EnDe as a community resource to advance safe deployment of MT in information retrieval and conversational assistants, particularly in emerging contexts such as wearable AI devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。