为古典文本翻译创建多参考标准,提升评估公平性。
PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarks
- 基于多个独立译本构建多参考评估框架。
- 包含1700段落、约34.5万词元的巴利文-英文基准数据集。
- 适用于佛教经典等需要多解释的古籍翻译评估。
数字人文项目越来越多依赖机器翻译和大语言模型,以扩大对古典、宗教及其它低翻译文本传统的影响。然而,现有翻译基准通常仅对比单一参考译文,而古典文本常允许多种忠实译法,差异体现在术语、语体和诠释上。本文提出PaliBench,不仅是一个针对巴利语到英语翻译的基准,也是一套可复用的多参考翻译基准构建方法。该案例基于《经藏》选段,结合了比丘苏贾托、比丘丹尼萨罗与比丘波提三位学者的独立英译。流程包括:大模型辅助的译文分段对齐、自动校验源文件、段级质量筛选、公式化重复去重以及多指标多参考评估。最终基准包含1,700段落、8,389个片段,约34.5万词元。我们用其评估了十款主流大语言模型,发现系统排名在不同指标间高度一致,但可靠性与语义异常率差异显著。核心贡献在于方法论:将既有学术译本转化为评估基础设施,不预设任何一译本为唯一标准。虽聚焦巴利佛典,该方法可推广至其他具备充分独立译本的古典文献体系。
原文摘要 · Abstract (English)
Digital humanities projects increasingly rely on machine translation and large language models to widen access to classical, religious, and otherwise under-translated textual traditions. Yet standard translation benchmarks are poorly suited to such materials: they typically compare a system output against a single reference translation, even though classical texts often support multiple faithful renderings that differ in terminology, register, and interpretation. This article introduces PaliBench, both a benchmark for Pali-to-English translation and a reusable method for constructing multi-reference translation benchmarks for classical languages. The Pali case study draws on passages from the Sutta Pitaka aligned with independent English translations by Bhikkhu Sujato, Bhikkhu Thanissaro, and Bhikkhu Bodhi. The workflow combines LLM-assisted alignment of independently segmented translations, automated verification against source files, passage-level quality filtering, deduplication of formulaic repetitions, and multi-metric evaluation against multiple human references. The resulting benchmark contains 1,700 passages spanning 8,389 segments and approximately 345,000 tokens. We use it to evaluate ten contemporary large language models with complementary metrics, finding strong cross-metric concordance in system rankings alongside substantial variation in reliability and semantic outlier rates. The broader contribution is methodological: PaliBench shows how existing scholarly translations can be transformed into evaluation infrastructure for interpretive textual traditions without treating any single translation as definitive. Although developed for Pali Buddhist texts, the approach could be portable to other classical corpora where sufficient independent reference translations exist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。