arXiv:2410.11672cs.CLcs.AI2024-10被引 8

简单词元可预测大模型测评答案,暗示评测存在漏洞。

Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers

  • 用简单词元构建分类器,无需理解即可高分通过评测
  • 多个现代测评中,词元模型得分接近大模型表现
  • 提醒研究者警惕评测结果被表面模式误导

AI评测的可靠性对准确评估系统能力至关重要。本文探讨评测内部有效性问题:模型是否可能通过非预期方式解题,绕过真正测试的能力。这种现象被称为‘聪明汉斯效应’,即利用表面线索而非真实推理。已有研究发现,早期NLP评测中,如‘not’等单个n-gram可高度预测标签,且监督模型会利用这些模式。本文研究现代多选题评测中,仅用简单n-gram能否预测答案,并考察大模型是否依赖此类模式。结果显示,仅基于n-gram的简单分类器在多个评测上达到高分,而无需具备被测能力。同时,证据表明现代大模型可能也在使用这些表面模式解题。这提示当前评测的内部有效性可能受损,解读大模型性能时需谨慎。

原文摘要 · Abstract (English)

The integrity of AI benchmarks is fundamental to accurately assess the capabilities of AI systems. The internal validity of these benchmarks - i.e., making sure they are free from confounding factors - is crucial for ensuring that they are measuring what they are designed to measure. In this paper, we explore a key issue related to internal validity: the possibility that AI systems can solve benchmarks in unintended ways, bypassing the capability being tested. This phenomenon, widely known in human and animal experiments, is often referred to as the 'Clever Hans' effect, where tasks are solved using spurious cues, often involving much simpler processes than those putatively assessed. Previous research suggests that language models can exhibit this behaviour as well. In several older Natural Language Processing (NLP) benchmarks, individual $n$-grams like "not" have been found to be highly predictive of the correct labels, and supervised NLP models have been shown to exploit these patterns. In this work, we investigate the extent to which simple $n$-grams extracted from benchmark instances can be combined to predict labels in modern multiple-choice benchmarks designed for LLMs, and whether LLMs might be using such $n$-gram patterns to solve these benchmarks. We show how simple classifiers trained on these $n$-grams can achieve high scores on several benchmarks, despite lacking the capabilities being tested. Additionally, we provide evidence that modern LLMs might be using these superficial patterns to solve benchmarks. This suggests that the internal validity of these benchmarks may be compromised and caution should be exercised when interpreting LLM performance results on them.

大模型评测词元分析评测漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。