arXiv:2509.20557cs.CL2025-09被引 1

构建首个中文方言机翻错误标注数据集,助力低资源语言翻译优化

SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages

  • 基于英译中/粤/吴语平行语料,标注错误区间、类型与严重程度
  • 多模型测试显示当前大模型错误检测精度有限,凸显数据价值
  • 专为翻译质量评估与错误感知生成设计,适合低资源语言研究者

尽管近年来机器翻译取得显著进展,但许多低资源语言仍因缺乏大规模训练数据和语言资源而受限。本文提出 \\(\text{SiniticMTError}\\) ,一个细粒度的新数据集,基于现有双语语料库,对英译普通话、粤语、吴语及由非平行源生成的普-闽南语翻译样本中的机器翻译错误进行错误区间、错误类型和错误严重程度标注。该数据集可支持翻译质量估计、错误感知生成与低资源语言评估等研究,助力模型微调以提升错误检测能力。我们还使用多种开源与闭源大模型,通过段落级与相关性指标进行基准测试,揭示其错误检测精度普遍较低,凸显本数据集的必要性。最后,我们报告了母语者主导的严谨标注流程,包括试点研究、迭代反馈分析,以及对错误类型与严重程度分布的深入洞察。

原文摘要 · Abstract (English)

Despite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic resources. In this paper, we introduce \dsname, a novel fine-grained dataset that builds on existing parallel corpora to provide error span, error type, and error severity annotations in machine-translated examples from English to Mandarin, Cantonese, and Wu Chinese, along with a Mandarin-Hokkien component derived from a non-parallel source. Our dataset serves as a resource for the MT community to fine-tune models with error detection capabilities, supporting research on translation quality estimation, error-aware generation, and low-resource language evaluation. We also establish baseline results using language models to benchmark translation error detection performance. Specifically, we evaluate multiple open source and closed source LLMs using span-level and correlation-based MQM metrics, revealing their limited precision, underscoring the need for our dataset. Finally, we report our rigorous annotation process by native speakers, with analyses on pilot studies, iterative feedback, insights, and patterns in error type and severity.

机器翻译错误标注低资源语言中文方言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。