arXiv:2608.21827cs.CL2026-08

首个评测大模型读懂现代汉语诗歌逻辑的基准,发现现有模型表现有限。

Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

论文配图:Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?
图 1 · 摘自论文原文
  • 构建诗歌逻辑四任务三层次评估框架,覆盖段、行、意象
  • 六款主流大模型在非思考与思考模式下均表现不佳
  • 为中文诗歌理解提供新评估标准,适合研究文学AI者参考

大型语言模型(LLMs)在众多自然语言处理任务中取得显著进展,但其对文学文本——特别是现代汉语诗歌——的理解能力仍鲜有研究。现代汉语诗歌的独特文学特征要求一种不同于常规文本的推理方式,其“诗意逻辑”需超越表面语义分析,进行整体性理解。然而,当前评估范式普遍忽略这一关键维度。为此,我们提出Peony,首个专为评估现代汉语诗歌诗意逻辑而设计的基准。我们将诗意逻辑定义为涵盖段、行、意象三个层级的四项任务,并基于Peony系统评估与分析六款主流大模型,对比其在非思考与思考配置下的表现。实验结果揭示了当前大模型在理解现代汉语诗歌诗意逻辑方面的局限性,并验证了Peony的有效性与必要性。数据与代码将公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved significant progress across a wide range of natural language processing (NLP) tasks, yet their ability to understand literary texts, particularly modern Chinese poetry, remains largely unexplored. The unique literary characteristics of modern Chinese poetry necessitate a distinct form of reasoning for effective comprehension. Unlike conventional texts that convey clear information, the unique "poetic logic" of modern Chinese poetry requires a holistic reasoning approach that goes beyond superficial semantic analysis to be understood. However, current evaluation paradigms largely ignore this critical dimension. To address this gap, we propose Peony, the first benchmark specifically designed for evaluating the poetic logic of modern Chinese poetry. We define poetic logic as four tasks across three levels, namely stanza, line, and imagery, and systematically evaluate and analyze six mainstream LLMs based on Peony. We evaluate these models under both non-thinking and thinking configurations. The experimental results reveal the limitations of current LLMs in understanding the poetic logic of modern Chinese poetry and validate the effectiveness and necessity of Peony. Our data and code will be available.

诗歌理解大模型评估诗意逻辑中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。