arXiv:2509.18401cs.CL2025-09EMNLP被引 5

评估大模型在波斯文学创作中的创造力,涵盖诗歌与短篇小说。

Evaluating the Creativity of LLMs in Persian Literary Text Generation

  • 基于托兰斯创造思维测试,从原创性、流畅性等四维度评估模型生成能力。
  • 采用大模型自动评分,与人工评判一致性高,有效降低评估成本。
  • 分析隐喻、拟人等四大修辞手法的运用,揭示模型优劣与改进方向。

大型语言模型在生成文学文本(包括诗歌和短篇小说)方面展现了显著的创造力。然而,以往研究主要集中于英语,对非英语文学传统关注不足,且缺乏标准化的创造力评估方法。本文评估大模型在生成富含文化相关表达的波斯文学文本方面的能力。我们构建了一个涵盖20个不同主题的用户生成波斯文学数据集,并通过改编托兰斯创造思维测试,从原创性、流畅性、灵活性和丰富性四个维度评估模型输出。为降低评估成本,采用大模型作为裁判进行自动化评分,并通过组内相关系数验证其与人工判断的一致性,结果显示高度一致。此外,我们分析了模型对四种核心修辞手法——明喻、隐喻、夸张和对照——的理解与运用能力。结果表明,大模型在波斯文学生成中既有优势也存在局限,凸显了进一步优化的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated notable creative abilities in generating literary texts, including poetry and short stories. However, prior research has primarily centered on English, with limited exploration of non-English literary traditions and without standardized methods for assessing creativity. In this paper, we evaluate the capacity of LLMs to generate Persian literary text enriched with culturally relevant expressions. We build a dataset of user-generated Persian literary spanning 20 diverse topics and assess model outputs along four creativity dimensions-originality, fluency, flexibility, and elaboration-by adapting the Torrance Tests of Creative Thinking. To reduce evaluation costs, we adopt an LLM as a judge for automated scoring and validate its reliability against human judgments using intraclass correlation coefficients, observing strong agreement. In addition, we analyze the models' ability to understand and employ four core literary devices: simile, metaphor, hyperbole, and antithesis. Our results highlight both the strengths and limitations of LLMs in Persian literary text generation, underscoring the need for further refinement.

文学生成波斯语创造力评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。