提升低资源语言民间故事的模型表示能力,增强文化数字化研究
Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages
- 用分类任务评估大模型对低资源语言的故事表征能力
- 长序列适配与民间故事域继续预训练显著提升分类效果
- 适合从事文化遗产数字化、低资源语言处理的研究者
民间故事是了解文明社会与文化的重要知识资源。数字民俗学研究依赖文本数据的抽象表征来自动化理解这些故事。尽管多个大语言模型声称可处理如爱尔兰语、盖尔语等低资源语言,本文通过两项分类任务考察其表征有效性,并提出三种改进策略。结果表明,将模型适配为支持更长序列输入,并在民间故事领域继续预训练,能有效提升分类性能。但这一提升被一个基于非上下文特征的基准SVM模型的优异表现所削弱。
原文摘要 · Abstract (English)
Folktales are a rich resource of knowledge about the society and culture of a civilisation. Digital folklore research aims to use automated techniques to better understand these folktales, and it relies on abstract representations of the textual data. Although a number of large language models (LLMs) claim to be able to represent low-resource langauges such as Irish and Gaelic, we present two classification tasks to explore how useful these representations are, and three adaptations to improve the performance of these models. We find that adapting the models to work with longer sequences, and continuing pre-training on the domain of folktales improves classification performance, although these findings are tempered by the impressive performance of a baseline SVM with non-contextual features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。