研究人类与GPT在摘要评价中依赖的特征,发现可提升GPT判断准确性的方法。
Exploring the features used for summary evaluation by Human and GPT
- 通过统计与机器学习分析,识别人类与GPT评价摘要时使用的共性特征。
- 指令GPT使用人类常用的评估指标,使其评分更接近人类判断。
- 为自动化摘要评估提供可解释性依据,适合模型评估与对齐研究者。
摘要评估需判断生成摘要是否准确反映原文核心思想与含义,要求深入理解内容。大语言模型(LLMs)被用于自动化该过程,充当评判者对摘要质量进行评分。尽管已有研究探讨了LLMs与人类评分的一致性,但尚不清楚它们在特定评价维度下具体利用了哪些属性或特征,也缺乏对评分与量化指标之间映射关系的关注。本文通过分析统计与机器学习指标,揭示了人类与生成式预训练模型(GPTs)评分所依赖的共性特征。进一步表明,引导GPT使用人类常用的评估标准,可提升其判断准确性,使其评分更贴近人类意见。
原文摘要 · Abstract (English)
Summary assessment involves evaluating how well a generated summary reflects the key ideas and meaning of the source text, requiring a deep understanding of the content. Large Language Models (LLMs) have been used to automate this process, acting as judges to evaluate summaries with respect to the original text. While previous research investigated the alignment between LLMs and Human responses, it is not yet well understood what properties or features are exploited by them when asked to evaluate based on a particular quality dimension, and there has not been much attention towards mapping between evaluation scores and metrics. In this paper, we address this issue and discover features aligned with Human and Generative Pre-trained Transformers (GPTs) responses by studying statistical and machine learning metrics. Furthermore, we show that instructing GPTs to employ metrics used by Human can improve their judgment and conforming them better with human responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。