arXiv:2606.01136cs.CL2026-06

用多参考译文评估佛典翻译,发现模型错误集中在高离群区域。

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication

  • 以多位译者构建参考集,用嵌入漂移筛选可疑翻译
  • 高漂移项中重大错误率超50%,但多数漂移项实为合理变体
  • GPT-5.5表现最优,Grok 4.3错误率最高且异常值最多

单分数翻译评估易将合理差异误判为错误,尤其在古典语言中更为显著。本文对四个主流大模型(GPT-5.5、Claude Sonnet 4.6、Gemini 3.1 Pro、Grok 4.3)在1700段巴利语经文上的英译进行审计,采用三位资深译者(Bhikkhu Sujato、Thanissaro Bhikkhu、Bhikkhu Bodhi)的译文作为局部参考包,而非单一标准答案。每个候选译文的归一化嵌入漂移距离参考中心点,作为初步筛查信号;超过1.5漂移阈值的1203个候选再由盲评三模型判断小组审定,该小组基于300例作者标注验证集校准。结果显示:漂移反映严重性而非错误本身——1.5-2.0漂移带内重大错误率仅7.9%,而高于3.0时升至51.6%;约80%的1.5-2.0漂移项被判定为有效变体。模型差异在高漂移尾部最明显:GPT-5.5的高漂移重大错误率最低,与Claude Sonnet 4.6和Gemini 3.1 Pro重叠;Grok 4.3则异常值最多,整体重大错误率27.6%,3.0以上达74.4%。主要错误类型(如遗漏、截断、教义术语错误)极易误导教义文本读者。本研究提出可复用的古典-现代翻译审计框架:定义多译者参考包,用嵌入漂移优先审查,仅对标记尾部进行审定,而非将离群视为错误。

原文摘要 · Abstract (English)

Single-score translation metrics can conflate legitimate variation with error, a problem especially acute for classical languages where multiple defensible English renderings of the same passage coexist. We audit Pali-to-English output from four flagship large language models (LLMs): GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro, and Grok 4.3, on 1,700 passages from the Pali Canon, using three established human translations by Bhikkhu Sujato, Thanissaro Bhikkhu, and Bhikkhu Bodhi as a local reference envelope rather than a single gold standard. Each candidate's normalized embedding drift from the reference centroid serves as a triage signal, not an error label; the 1,203 candidates above a 1.5 drift threshold are then adjudicated by a blinded three-model LLM judge panel, calibrated against a 300-instance author-adjudicated validation set. Two results stand out. First, drift predicts severity rather than error per se: the major-error rate among adjudicated high-drift candidates rose monotonically from 7.9% in the 1.5-2.0 band to 51.6% above 3.0, while approximately 80% of 1.5-2.0 outliers were judged valid translation variations. Second, model differences were clearest in the high-drift tail: GPT-5.5 had the lowest adjudicated high-drift major-error rate, with confidence intervals overlapping those of Claude Sonnet 4.6 and Gemini 3.1 Pro; Grok 4.3 had both the largest outlier volume and the highest tail major-error rate (27.6% overall, 74.4% above drift 3.0). The dominant major-error categories (e.g. omission or truncation, doctrinal term errors) are precisely the failures most likely to mislead readers of doctrinal text. The contribution is a reusable audit design for classical-to-modern translation: define a local reference envelope from multiple human translators, use embedding drift to prioritize review, and adjudicate the flagged tail rather than treating outlier status as error.

翻译审计古典语言多参考大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。