测试发现大模型在细微改动下准确率暴跌,暴露其依赖表面线索的弱点。
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
- 通过控制变量的扰动测试模型泛化能力
- 选项长度变化导致最高30分降幅,如Qwen 2.5 1.5B从89降至36
- 适合评估模型鲁棒性或研究偏差来源的研究者
本文提出一种「泛化压力测试」,评估大语言模型在选项长度、问题类型及无关名词替换等轻微且可控扰动下的泛化能力。尽管基准测试表现优异,模型在这些仅改变格式或词汇但不改内容的微小调整下,准确率急剧下降并出现意外偏差(如偏好更长干扰项)。例如,Qwen 2.5 1.5B在选项长度变化时,MMLU得分从60升至89后又骤降至36;GPT4o在问题类型改变时损失25分,三类扰动下平均下降6分。结果表明,大模型严重依赖表面线索,难以形成跨格式、词汇变化和无关内容转移的稳健抽象表征。
原文摘要 · Abstract (English)
In this paper, we propose a ``Generalization Stress Test" to assess Large Language Models' (LLMs) generalization ability under slight and controlled perturbations, including option length, problem types, and irrelevant noun replacements. We achieve novel and significant findings that, despite high benchmark scores, LLMs exhibit severe accuracy drops and unexpected biases (e.g., preference for longer distractors) when faced with these minor but content-preserving modifications. For example, Qwen 2.5 1.5B's MMLU score rises from 60 to 89 and drops from 89 to 36 when option lengths are changed without altering the question. Even GPT4o experiences a 25-point accuracy loss when problem types are changed, with a 6-point drop across all three modification categories. These analyses suggest that LLMs rely heavily on superficial cues rather than forming robust, abstract representations that generalize across formats, lexical variations, and irrelevant content shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。