arXiv:2602.24119cs.CLcs.AI2026-02被引 1

评估大模型翻译古希腊医学哲学文本,发现术语罕见是失败主因。

Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen

  • 用专家无参考评估法,测试三款大模型对盖伦文本的翻译质量。
  • 术语越罕见,翻译越差,极端情况甚至完全失败,整体得分79.9分。
  • 首次揭示术语频率是预测翻译失败的关键因素,适合古籍数字化研究者。

本研究评估商业大语言模型(LLM)在翻译古希腊语技术性散文方面的表现,并将自动化评估指标与领域专家的人工判断进行对比。实验选取公元2世纪希腊医师盖伦的两部作品:一篇有已知英文译本的阐述性文本和一部从未被翻译过的药理学文本,共20段落,由ChatGPT、Claude、Gemini三款模型生成60次翻译。质量评估采用七种自动化指标,并通过改进的多维质量度量(MQM)框架进行无参考式专家人工评价。结果显示,在有参考的阐述性文本上,模型平均得分95.2/100;而在未翻译的药理学文本上,平均分降至79.9/100,且呈双峰分布:两个术语密度极高的段落出现灾难性失败,其余段落得分仅比阐述性文本低4分以内。术语稀有度(以语料库频次衡量)成为预测失败的主导因素(相关系数r = -0.97)。自动化指标仅在质量差异大的文本中与人工判断有中等相关性,无法区分高质量翻译。本研究为首个针对任何古代语言的系统性、无参考式专家评估,也是首个识别出预测翻译失败文本特征的研究。

原文摘要 · Abstract (English)

Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and benchmarks standard automated MT evaluation metrics against expert human judgment. Design: We evaluated 60 translations by three LLMs (ChatGPT, Claude, Gemini) of 20 paragraph-length passages from 2 works by the Greek physician Galen (c. 129-216 CE): an expository text with two published English translations and a pharmacological text never before translated. Quality was assessed using seven automated metrics and systematic reference-free human evaluation via a modified Multidimensional Quality Metrics (MQM) framework applied by domain specialists. Findings: On the translated expository text, LLMs achieved high quality (mean MQM score 95.2/100). On the untranslated pharmacological text, quality was lower (79.9/100) but bimodally distributed: two passages with extreme terminological density produced catastrophic failures, while remaining passages scored within 4 points of the expository text. Terminology rarity, operationalized via corpus frequency, emerged as the dominant predictor of failure (r = -.97). Automated metrics showed moderate correlation with human judgment only on texts with wide quality variance; no metric discriminated among high-quality translations. Originality: This is the first systematic, reference-free expert human evaluation of LLM translation for any ancient language and the first study identifying textual properties predictive of translation failure.

大模型翻译古希腊语术语稀有度医疗文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。