现有机器翻译质量评估方法无法独立使用,因忽略文本连贯性等关键语言特征。
Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

- 从段落级评估翻译质量,忽视连贯性与修辞特征。
- 存在泛化失败、数据稀缺、标注困难等结构性缺陷。
- 适合关注翻译质量管理的从业者,尤其需警惕自动化误判风险。
自动化翻译质量评估(QE)被广泛视为规模化管理翻译质量的途径,但相关研究和测试常缺乏透明度与可复现性。本文从理论与实证角度分析了当前QE技术的根本局限:孤立段落级别的质量评估难以捕捉语篇连贯性、逻辑一致性及风格特征。实证研究揭示多重问题,包括泛化能力差、系统性偏差、过拟合与分布坍塌、性能差距、错误标注难题及数据匮乏。这些缺陷源于人类语言与翻译作为认知行为的复杂性,单纯增加数据或优化架构尚未解决。因此,段落级QE评分不应作为生产中路由、发布或跳过人工审校的唯一依据;未来应聚焦基于MQM标准的自动化人工评估。
原文摘要 · Abstract (English)
Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing number of tools and technologies have been released in pursuit of this goal. However, the proliferation of new QE systems has not always been accompanied by robust, transparent, and reproducible research and testing. This gap deserves critical scrutiny. This paper examines some fundamental limitations of the QE technology from both theoretical and empirical perspectives, arguing that current QE systems are structurally ill-equipped to serve as reliable standalone tools in real-world translation workflows. The reviewed evidence suggests that QE suffers from a range of interrelated and largely unresolved limitations. Most fundamentally, the evaluation of the quality of translation at the level of isolated segments is problematic because it tends to miss out on cohesion, coherence, and stylistic and rhetorical text features. In addition, empirical research documents several other limitations and flaws, including failure to generalize, systematic biases, overfitting and distribution collapse, performance gaps, error annotation challenges, and data scarcity. These are structural limitations arising from the complexity of human language and translation as a cognitive and communicative act - limitations that more data and better architectures have so far not overcome. Consequently, segment-level QE scores should not be used as a standalone basis for routing, release, or review bypass in production; we argue future work should focus on automating human evaluation grounded in MQM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。