测试大模型对日本妖怪传说的认知,发现日语训练模型表现更优。
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
- 构建809道关于日本妖怪的多选题评测集,评估模型文化知识。
- 基于日语持续预训练的Llama-3模型准确率最高,显著优于英语主导模型。
- 为非英语文化认知研究提供可复用的数据与基准,适合文化计算研究者。
尽管大型语言模型在多种语言中展现出强大的语言理解与生成能力,但其文化知识常局限于英语社区,容易边缘化非英语文化。为解决此问题,已有研究探讨了模型的文化意识评估及提升方法。本研究聚焦于民间故事这一文化传播核心载体,特别关注日本民间故事中的妖怪(Yokai)。妖怪是源自日本民间传说的超自然生物,至今仍广泛出现在艺术与娱乐作品中,是文化表达的重要媒介。为此,我们构建了YokaiEval——一个包含809个多项选择题(每题四个选项)的基准数据集,用于考察模型对妖怪的知识掌握情况。我们评估了31个日语及多语言大模型在此数据集上的表现。结果表明,经过日语资源训练的模型准确率显著高于以英语为中心的模型,其中采用日语持续预训练的Llama-3系列模型表现尤为突出。相关代码与数据集已公开于https://github.com/CyberAgentAILab/YokaiEval。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) have demonstrated strong language understanding and generation abilities across various languages, their cultural knowledge is often limited to English-speaking communities, which can marginalize the cultures of non-English communities. To address the problem, evaluation of the cultural awareness of the LLMs and the methods to develop culturally aware LLMs have been investigated. In this study, we focus on evaluating knowledge of folktales, a key medium for conveying and circulating culture. In particular, we focus on Japanese folktales, specifically on knowledge of Yokai. Yokai are supernatural creatures originating from Japanese folktales that continue to be popular motifs in art and entertainment today. Yokai have long served as a medium for cultural expression, making them an ideal subject for assessing the cultural awareness of LLMs. We introduce YokaiEval, a benchmark dataset consisting of 809 multiple-choice questions (each with four options) designed to probe knowledge about yokai. We evaluate the performance of 31 Japanese and multilingual LLMs on this dataset. The results show that models trained with Japanese language resources achieve higher accuracy than English-centric models, with those that underwent continued pretraining in Japanese, particularly those based on Llama-3, performing especially well. The code and dataset are available at https://github.com/CyberAgentA ILab/YokaiEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。