提出非线性评分模型,让翻译质量评估更符合人类感知习惯。
Non-Linear Scoring Model for Translation Quality Evaluation
- 用对数函数建模错误容忍度,随样本量增长而递增但增速放缓。
- 实证发现:错误可接受数量与样本长度呈对数关系,非线性增长。
- 适用于人工与AI生成翻译,提升评估公平性与一致性。
基于多维质量度量(MQM)的分析型翻译质量评估(TQE)传统上采用线性误差-扣分映射,基于1000-2000词参考样本校准。然而,线性外推会导致不同长度样本的判断偏差:短样本被过度惩罚,长样本则惩罚不足,与专家直觉不符。本文基于多范围框架,提出一种校准的非线性评分模型,更准确反映人类在不同样本长度下的质量感知。三个大规模企业环境的实证数据表明,可接受错误数随样本量对数增长,而非线性。心理物理学和认知负荷理论(如韦伯-费希纳定律)支持此假设,解释了为何额外错误的感知影响随规模增大而减弱,而认知负担持续上升。本文提出双参数模型 E(x) = a * ln(1 + b * x),a, b > 0,以参考容差为锚点,通过两点校准与一维根查找实现动态调整。该模型定义了线性近似保持在±20%相对误差内的区间,并仅需替换原有容差函数即可集成至现有评估流程。方法显著提升人类与AI生成翻译评估的可解释性、公平性与评分者间一致性。通过实现感知有效的评分范式,推动翻译质量评估向更精确、可扩展的方向发展。该模型也为基于AI的文档级评估提供了与人类判断对齐的更强基础。文章讨论了对CAT/LQA系统的实施考量及对人机生成文本评估的影响。
原文摘要 · Abstract (English)
Analytic Translation Quality Evaluation (TQE), based on Multidimensional Quality Metrics (MQM), traditionally uses a linear error-to-penalty scale calibrated to a reference sample of 1000-2000 words. However, linear extrapolation biases judgment on samples of different sizes, over-penalizing short samples and under-penalizing long ones, producing misalignment with expert intuition. Building on the Multi-Range framework, this paper presents a calibrated, non-linear scoring model that better reflects how human content consumers perceive translation quality across samples of varying length. Empirical data from three large-scale enterprise environments shows that acceptable error counts grow logarithmically, not linearly, with sample size. Psychophysical and cognitive evidence, including the Weber-Fechner law and Cognitive Load Theory, supports this premise by explaining why the perceptual impact of additional errors diminishes while the cognitive burden grows with scale. We propose a two-parameter model E(x) = a * ln(1 + b * x), a, b > 0, anchored to a reference tolerance and calibrated from two tolerance points using a one-dimensional root-finding step. The model yields an explicit interval within which the linear approximation stays within +/-20 percent relative error and integrates into existing evaluation workflows with only a dynamic tolerance function added. The approach improves interpretability, fairness, and inter-rater reliability across both human and AI-generated translations. By operationalizing a perceptually valid scoring paradigm, it advances translation quality evaluation toward more accurate and scalable assessment. The model also provides a stronger basis for AI-based document-level evaluation aligned with human judgment. Implementation considerations for CAT/LQA systems and implications for human and AI-generated text evaluation are discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。