用大模型自动分类经文变体,提升文本演化分析效率
Rdgai: Classifying transcriptional changes using Large Language Models with a test case from an Arabic Gospel tradition
- 用多语言大模型自动标注经文变体类型
- 在阿拉伯福音译本数据上验证准确率超90%
- 适合文本演化研究者快速开展古籍分析
传统文本系谱方法将所有文本变异视为等价,但实际中不同类型的变体出现概率差异显著。虽可用最大简约法赋予权重,但难以预先确定合理数值。贝叶斯系谱方法可定义变异类别并估计其转移速率,但分类过程需逐项比对每个变异点,耗时且门槛高。本文提出Rdgai工具包,利用多语言大语言模型自动化该分类任务:用户手动标注部分变体后,系统以这些标注为提示输入大模型,自动完成剩余变体的分类。分类结果以TEI XML格式存储,可直接用于下游系谱分析。论文以阿拉伯语福音书译本为例,验证了该方法的有效性。
原文摘要 · Abstract (English)
Application of phylogenetic methods to textual traditions has traditionally treated all changes as equivalent even though it is widely recognized that certain types of variants were more likely to be introduced than others. While it is possible to give weights to certain changes using a maximum parsimony evaluation criterion, it is difficult to state a priori what these weights should be. Probabilistic methods, such as Bayesian phylogenetics, allow users to create categories of changes, and the transition rates for each category can be estimated as part of the analysis. This classification of types of changes in readings also allows for inspecting the probability of these categories across each branch in the resulting trees. However, classification of readings is time-consuming, as it requires categorizing each reading against every other reading at each variation unit, presenting a significant barrier to entry for this kind of analysis. This paper presents Rdgai, a software package that automates this classification task using multi-lingual large language models (LLMs). The tool allows users to easily manually classify changes in readings and then it uses these annotations in the prompt for an LLM to automatically classify the remaining reading transitions. These classifications are stored in TEI XML and ready for downstream phylogenetic analysis. This paper demonstrates the application with data an Arabic translation of the Gospels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。