用大模型模拟诗人视角,提升古诗理解评估准确性
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?

- 让大模型扮演诗人角色,从创作视角评估诗歌理解
- 在修辞手法和陌生化维度上误差降低超90%
- 适合需要高质量诗歌评估的研究者与评测系统
传统自动评估方法因古典汉语诗歌的独特性而效果不佳。人工评估虽可靠但成本高,难以大规模应用。本文提出Poller(Poetry LLM Evaluator),一种利用大语言模型(LLMs)评估诗歌理解任务的新方法。具体而言,该方法让LLM以诗人身份,基于详细背景信息进行判断,从而模拟人类评价。我们在多个大模型上进行了实验,对诗歌解读在八个专业维度上的表现进行评估。结果表明,该方法显著降低了大模型与人类之间的评估误差。尤其在修辞手法和陌生化两个维度上,相比基线方法分别实现94.55%和89.53%的误差减少。这些性能是传统大模型评估方法无法达到的。多模型、多维度的实验验证了该方法的有效性。本工作弥合了自动化效率与人类专家判断之间的差距,为诗歌相关任务的自动化评估奠定了基础。
原文摘要 · Abstract (English)
Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry because of the distinct nature of this literary genre. Human evaluation remains reliable, but is expensive and not applicable to large-scale data. In this paper, we propose Poller (Poetry LLM Evaluator), a novel method leveraging large language models (LLMs) to evaluate the poetry understanding task. Specifically, our method requires LLMs to play the role of a poem's author with detailed information, thereby emulating human evaluation and judgment by adopting the poet's perspective. We conducted comprehensive experiments on multiple LLMs, evaluating the interpretations of poems across eight specialized dimensions. Experimental results demonstrate that our method effectively reduces the evaluation error between LLMs and humans. Especially for specific dimension evaluation, Poller-based LLMs achieve a 94.55% and 89.53% error reduction for rhetorical techniques and defamiliarization, respectively, compared to baseline methods. These performances are unattainable by conventional LLM evaluation methods. Experimental results from multiple LLMs across various dimensions validate the efficacy of our method. This work bridges the gap between automated efficiency and human expertise, establishing a foundation for automated evaluation in poetry-related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。