用大模型解析古诗,探索文学理解的规律。
Understanding Literary Texts by LLMs: A Case Study of Ancient Chinese Poetry
- 基于大模型设计诗歌理解指标,量化评估古诗
- 发现不同诗集间存在可识别的文学模式
- 为高阶文学生成提供技术支撑,适合研究者参考
大语言模型(LLMs)的兴起在文学领域引发广泛关注。曾被认为难以实现的AI文学创作正逐步成为现实,在诗歌、笑话和短篇故事等体裁中已涌现出诸多工具,带来全新视角。然而,作品质量提升面临瓶颈,主要源于文学理解与欣赏需具备较高的知识门槛,如文学理论、审美感知及跨学科知识,而该领域权威数据严重不足。此外,文学作品评价复杂且难于完全量化,直接制约了AI创作的进一步发展。为此,本文尝试从大模型视角探索文学文本的奥秘,以古代汉语诗歌为例进行实验。首先,从多个来源收集多种古诗,并由专家标注其中一部分;其次,设计一系列基于大模型的诗歌理解度量指标,用于评估全部诗歌;最后,分析不同诗集间的相关性与差异性,识别文学模式。实验中观察到一系列启发性现象,为未来基于大模型的高级文学创作提供了技术依据。
原文摘要 · Abstract (English)
The birth and rapid development of large language models (LLMs) have caused quite a stir in the field of literature. Once considered unattainable, AI's role in literary creation is increasingly becoming a reality. In genres such as poetry, jokes, and short stories, numerous AI tools have emerged, offering refreshing new perspectives. However, it's difficult to further improve the quality of these works. This is primarily because understanding and appreciating a good literary work involves a considerable threshold, such as knowledge of literary theory, aesthetic sensibility, interdisciplinary knowledge. Therefore, authoritative data in this area is quite lacking. Additionally, evaluating literary works is often complex and hard to fully quantify, which directly hinders the further development of AI creation. To address this issue, this paper attempts to explore the mysteries of literary texts from the perspective of LLMs, using ancient Chinese poetry as an example for experimentation. First, we collected a variety of ancient poems from different sources and had experts annotate a small portion of them. Then, we designed a range of comprehension metrics based on LLMs to evaluate all these poems. Finally, we analyzed the correlations and differences between various poem collections to identify literary patterns. Through our experiments, we observed a series of enlightening phenomena that provide technical support for the future development of high-level literary creation based on LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。