用大模型自动分析灰姑娘故事的母题,发现跨语言共性与差异。
Large language models for folktale type automation based on motifs: Cinderella case study
- 基于大模型识别灰姑娘变体中的叙事母题
- 通过聚类与降维揭示母题组合的相似与差异模式
- 适合对民间故事跨文化研究感兴趣的学者
人工智能正被广泛应用于数字人文等领域。本文构建了一套大规模民间故事分析方法,利用机器学习与自然语言处理技术,自动识别大量灰姑娘故事变体中的叙事母题,并通过聚类与降维分析其异同。结果表明,大语言模型能够捕捉故事中复杂的母题互动,实现对海量文本的计算分析,支持跨语言比较研究。
原文摘要 · Abstract (English)
Artificial intelligence approaches are being adapted to many research areas, including digital humanities. We built a methodology for large-scale analyses in folkloristics. Using machine learning and natural language processing, we automatically detected motifs in a large collection of Cinderella variants and analysed their similarities and differences with clustering and dimensionality reduction. The results show that large language models detect complex interactions in tales, enabling computational analysis of extensive text collections and facilitating cross-lingual comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。