arXiv:2510.08596cs.CL2025-10被引 1

提出新指标提升大模型创意文本评估准确性

Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs

  • 用输出概率分布构建信心得分,减少对创意文本的偏见
  • 在99个创意任务中,新指标使新颖回复偏好率从0%提升至19%
  • 能有效区分任务难度,适合评估现代大模型创造力

自参考指标如自困惑度对创造性文本生成存在严重偏差。本文提出基于模型输出概率分布的信心得分(CS),作为更少偏差的替代方案。在 gpt-4o-mini 上的实验表明,尽管基于流畅性的指标在99个创意提示中对新颖回答的偏好率为0%,而本方法可达19%,差异具有统计显著性(95%置信区间:[11.1%, 27.3%])。此外,信心得分可有效区分简单、中等和困难任务,各组置信区间无重叠。该指标在缓解传统指标的创意偏差的同时,保留其核心评估优势,为现代大模型提供更均衡的评估方式。

原文摘要 · Abstract (English)

Reference-free metrics like self-perplexity are strongly biased against creative text generation. We propose the Confidence Score (CS), derived from a model's output probability distribution, as a less biased alternative. Experiments on gpt-4o-mini show that while fluency-based metrics prefer novel responses in 0\% of cases on 99 creative prompts, our CS does so 19% of the time, a statistically significant difference (95% CI for difference: [11.1%, 27.3%]). We also show that CS effectively distinguishes between easy, medium, and hard tasks, confirmed by non-overlapping confidence intervals. The Confidence Score thus mitigates the creativity bias of traditional metrics while retaining their core evaluative strengths, offering a more balanced assessment for modern LLMs.

大模型评估创意生成指标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。