arXiv:2510.15313cs.CL2025-10被引 2

评测大模型生成唐诗的能力,发现其自评存在严重偏差。

Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry

  • 构建三步评估框架,融合计算指标、模型评分与专家评审
  • 六款主流大模型生成的唐诗普遍因格式错误被专家否定
  • 适合关注中文创作质量评估或文化复杂任务研究者

大型语言模型在创意领域应用日益广泛,但在古典汉诗生成与评价方面表现仍不清晰。本文提出一个包含计算指标、大模型作为评判者及人类专家验证的三步评估框架,对六款先进大模型在唐诗生成中的主题、情感、意象、格律与风格等多个维度进行评估。分析揭示显著的‘回音室’效应:大模型倾向于高估模仿统计模式但违反严格韵律规则的机器生成诗歌,与人类专家判断存在显著差异。研究强调了将大模型单独用于文化复杂任务评价的局限性,凸显混合人机验证框架的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to creative domains, yet their performance in classical Chinese poetry generation and evaluation remains poorly understood. We propose a three-step evaluation framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation. Using this framework, we evaluate six state-of-the-art LLMs across multiple dimensions of poetic quality, including themes, emotions, imagery, form, and style, in the context of Tang poetry generation. Our analysis reveals a critical "echo chamber" effect: LLMs systematically overrate machine-generated poems that mimic statistical patterns yet fail strict prosodic rules, diverging significantly from human expert judgments. These findings underscore the limitations of using LLMs as standalone evaluators for culturally complex tasks, highlighting the necessity of hybrid human-model validation frameworks.

大模型评估唐诗生成人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。