arXiv:2604.03695cs.CL2026-04被引 2

首个诗歌评估框架揭示大模型仍难企及人类诗人的创造力与情感表达。

POEMetric: The Last Stanza of Humanity

  • 构建多维度诗歌评估体系,涵盖形式、创意与整体质量。
  • 人类在创意、意象运用和情感共鸣上显著优于大模型,得分超3.9。
  • 适合研究生成式AI、诗歌创作或人机创造力对比的学者参考。

大型语言模型能创作诗歌,但距离人类诗人还有多远?本文提出首个综合性诗歌评估框架POEMetric,从三个方面进行评测:1)按特定形式与主题生成诗歌的基本指令遵循能力;2)创造力、词汇多样性、独特性、情感共鸣、意象与修辞手法等高级能力;3)整体诗歌质量与作者归属判断。研究团队构建了包含203首英文诗歌的人类诗集数据集,涵盖7种固定诗体,标注了格律、押韵模式与主题,并基于相同诗体与主题对30个大模型进行测试,共生成6,090篇模型诗歌。通过规则评估与大模型作为裁判两种方式,结果经人类专家验证。结果显示,尽管顶级模型在形式准确度(4.26/5.00)和主题契合度(4.99)上表现优异,但在创造性(4.02)、独特性(3.95)、情感共鸣(4.06)、意象使用(4.49)和修辞技巧(4.67)方面均落后于人类诗人;人类在整体诗歌质量评分上也远超最佳模型(4.22 vs. 3.20)。因此,诗歌生成仍是大模型面临的严峻挑战。数据与代码已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can compose poetry, but how far are they from human poets? In this paper, we introduce POEMetric, the first comprehensive framework for poetry evaluation, examining 1) basic instruction-following abilities in generating poems according to a certain form and theme, 2) advanced abilities of showing creativity, lexical diversity, and idiosyncrasy, evoking emotional resonance, and using imagery and literary devices, and 3) general appraisal of the overall poem quality and estimation of authorship. We curated a human poem dataset - 203 English poems of 7 fixed forms annotated with meter, rhyme patterns and themes - and experimented with 30 LLMs for poetry generation based on the same forms and themes of the human data, totaling 6,090 LLM poems. Based on POEMetric, we assessed the performance of both human poets and LLMs through rule-based evaluation and LLM-as-a-judge, whose results were validated by human experts. Results show that, though the top model achieved high form accuracy (4.26 out of 5.00, with Gemini-2.5-Pro as a judge; same below) and theme alignment (4.99), all models failed to reach the same level of advanced abilities as human poets, who achieved unparalleled creativity (4.02), idiosyncrasy (3.95), emotional resonance (4.06), and skillful use of imagery (4.49) and literary devices (4.67). Humans also defeated the best-performing LLM in overall poem quality (4.22 vs. 3.20). As such, poetry generation remains a formidable challenge for LLMs. Data and codes are released at https://github.com/Bingru-Li/POEMetric.

诗歌生成大模型评估创造力评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。