大模型用对比评估法,诗歌评分比外行人类更准。
Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
- 用小批量随机对比排序评估诗歌,避免主观偏差。
- Claude-3-Opus相关性达0.87,远超人类非专家的0.38。
- 适合需要客观艺术评价的场景,如创作辅助或筛选。
本研究将共识评估技术(CAT)应用于大型语言模型(LLMs),提出一种新的诗歌评价方法。基于90首诗歌的数据集,以发表平台作为真实基准,我们证明该方法使LLM显著优于非专家人类评委。通过在小型随机批次中进行强制选择排名,Claude-3-Opus与真实基准的斯皮尔曼等级相关系数达到0.87,远高于表现最佳的人类非专家(SRC=0.38)。LLM评估还展现出高一致性,表明该方法具有强鲁棒性。研究结果表明,在比较框架引导下,大模型可成为诗歌评价的有效且可靠的工具,为未来在其他创造性领域的应用铺平道路。
原文摘要 · Abstract (English)
This study adapts the Consensual Assessment Technique (CAT) for Large Language Models (LLMs), introducing a novel methodology for poetry evaluation. Using a 90-poem dataset with a ground truth based on publication venue, we demonstrate that this approach allows LLMs to significantly surpass the performance of non-expert human judges. Our method, which leverages forced-choice ranking within small, randomized batches, enabled Claude-3-Opus to achieve a Spearman's Rank Correlation of 0.87 with the ground truth, dramatically outperforming the best human non-expert evaluation (SRC = 0.38). The LLM assessments also exhibited high inter-rater reliability, underscoring the methodology's robustness. These findings establish that LLMs, when guided by a comparative framework, can be effective and reliable tools for assessing poetry, paving the way for their broader application in other creative domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。