构建首个多语言新约译本深度语料库,支持多层次翻译研究。
Targum -- A Multilingual New Testament Translation Corpus
- 整合12个在线圣经库与1个现有语料,收录651个新约译本
- 每种语言译本数达此前最丰富语料的2.4至5.0倍,如英语有194个独立版本
- 提供标准化元数据,支持按版本谱系或去重进行灵活分析
许多欧洲语言拥有丰富的圣经翻译历史,但现有语料库在追求语言广度时往往忽视深度。为弥补这一缺口,我们构建了一个包含651个新约译本的多语言语料库,其中334个为独立版本,涵盖英语(194个独立版本,共390个)、法语(41个,共78个)、意大利语(17个,共33个)、波兰语(29个,共48个)和西班牙语(53个,共102个),每种语言的译本数量是此前最丰富语料的2.4至5.0倍。语料来自12个在线圣经图书馆和一个已有语料库,每个译本均附带元数据,可映射到标准化标识符,包括作品、版本及修订年份。该标准化处理使研究者可根据需求定义‘唯一性’:既可对特定谱系(如KJV系)做微观分析,也可通过去重实现宏观研究。该语料库是首个在单语言深度上足以支撑多层级分析的多语言资源,填补了翻译史量化研究的空白。
原文摘要 · Abstract (English)
Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament translations, of which 334 are unique, spanning five languages with 2.4-5.0x more translations per language than any prior corpus: English (194 unique versions from 390 total), French (41 from 78), Italian (17 from 33), Polish (29 from 48), and Spanish (53 from 102). Aggregated from 12 online biblical libraries and one preexisting corpus, each translation is annotated with metadata that maps the text to a standardized identifier for the work, its specific edition, and its year of revision. This canonicalization allows researchers to define "uniqueness" for their own needs: they can perform micro-level analyses on translation families, such as the KJV lineage, or conduct macro-level studies by deduplicating closely related texts. By providing the first multilingual resource with sufficient depth per language for flexible, multilevel analysis, the corpus fills a gap in the quantitative study of translation history.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。