arXiv:2506.00514cs.CL2025-06ACL被引 9

评估常识生成的多样性评价指标,发现内容类指标更可靠。

Evaluating the Evaluation of Diversity in Commonsense Generation

  • 用大模型构建标注多样性数据集,系统评估现有评价指标。
  • 基于形式的指标会高估多样性,随机句也得分很高。
  • 内容类指标与人工评分高度相关,推荐未来研究使用。

在常识生成任务中,模型需在符合常识的基础上生成多视角回答。以往研究提出了大量基于形式与内容重叠的多样性评价指标,但其有效性尚不明确。本文开展系统性元评估,发现基于形式的指标普遍高估句子集合的多样性,甚至对随机生成的句子也给出过高分数。为此,我们利用大语言模型构建了一个新的、针对常识生成多样性进行标注的数据集,并在此基础上对现有指标进行元评估。实验表明,内容类指标显著优于形式类指标,与大模型的人工评分高度相关。因此建议未来常识生成研究应优先采用内容类指标评估输出多样性。

原文摘要 · Abstract (English)

In commonsense generation, given a set of input concepts, a model must generate a response that is not only commonsense bearing, but also capturing multiple diverse viewpoints. Numerous evaluation metrics based on form- and content-level overlap have been proposed in prior work for evaluating the diversity of a commonsense generation model. However, it remains unclear as to which metrics are best suited for evaluating the diversity in commonsense generation. To address this gap, we conduct a systematic meta-evaluation of diversity metrics for commonsense generation. We find that form-based diversity metrics tend to consistently overestimate the diversity in sentence sets, where even randomly generated sentences are assigned overly high diversity scores. We then use an Large Language Model (LLM) to create a novel dataset annotated for the diversity of sentences generated for a commonsense generation task, and use it to conduct a meta-evaluation of the existing diversity evaluation metrics. Our experimental results show that content-based diversity evaluation metrics consistently outperform the form-based counterparts, showing high correlations with the LLM-based ratings. We recommend that future work on commonsense generation should use content-based metrics for evaluating the diversity of their outputs.

常识生成多样性评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。