arXiv:2502.18798cs.CLcs.AI2025-02被引 2

新评分法能更真实评估大模型理解力,避免选项干扰。

Choices Speak Louder than Questions

  • 提出NPSQ评分法,分离问题与选项影响。
  • 传统方法受选项变化影响大,NPSQ则保持稳定。
  • 适合评估大模型真实理解能力的研究者使用。

近期研究质疑多选题问答(MCQA)评估是否真实反映大语言模型的理解能力。本文探讨了‘选项敏感性’——即模型决策更受答案选项影响而非真正理解问题的现象。提出一种新评分方法:基于问题的归一化概率偏移(NPSQ),旨在剥离选项干扰,更准确衡量理解力。在填空、符号及混合格式等多种输入形式下实验发现,传统基于对数似然或其长度归一化版本的方法易受选项表面特征影响;而NPSQ在选项修改后仍保持稳定,展现出更强的可靠性。

原文摘要 · Abstract (English)

Recent findings raise concerns about whether the evaluation of Multiple-Choice Question Answering (MCQA) accurately reflects the comprehension abilities of large language models. This paper explores the concept of choice sensitivity, which refers to the tendency for model decisions to be more influenced by the answer options than by a genuine understanding of the question. We introduce a new scoring method called Normalized Probability Shift by the Question (NPSQ), designed to isolate the impact of the question itself and provide a more reliable assessment of comprehension. Through experiments involving various input formats, including cloze, symbols, and hybrid formats, we find that traditional scoring methods - such as those based on log-likelihood or its length-normalized variant - are vulnerable to superficial characteristics of the answer choices. In contrast, NPSQ remains stable even when modifications are made to the answer options.

大模型评估问答系统评分方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。