大模型能理解规则背后的意图,而非仅依赖文字表面。
Evidence of conceptual mastery in the application of rules by Large Language Models
- 通过多组实验测试模型对规则的泛化应用能力
- 小模型更接近人类判断,且对提示敏感度更高
- 模型不依赖额外思考时间,但能把握规则意图
背景:大语言模型(LLMs)在生成类似人类判断时,并不一定体现真正的概念掌握,可能源于记忆或对任务细节的敏感。目标:在五个实验中,检验13个主流大模型在规则应用上的通用能力,包括文本与目的相悖的情形。方法:研究1A对比模型与新收集的人类数据在公开刺激和训练截止后生成的剧本上的表现;研究2A/2B引入限时指令,阻断人类判断的机制路径;研究3改变推理投入程度,模拟人类时间受限判断。结果:模型在两组刺激集上均紧密跟随人类判断,且在新刺激集中表现出与人类一致的“目的导向”倾向。对文本和目的的敏感性在提示变化下依然稳健。限时指令响应呈现模型特异性,表明概念能力与人类对齐存在差异。小参数模型更显著复现人类模式,且易受提示影响。增加推理努力未显著改变多数模型的规则应用,但GPT-oss与Claude Sonnet 5在高努力下仍显现明显目的导向趋势。尽管每模型温度校准以匹配人类方差,模型响应方差仍低于人类。结论:整体表明,大模型规则应用体现一种不依赖额外推敲的稳定语义理解能力。
原文摘要 · Abstract (English)
Background. Evidence that large language models (LLMs) reproduce human judgments does not establish conceptual mastery: the correspondence may reflect memorisation or be sensitivite to incidental task features. Objective. Across five experiments, we test whether 13 LLMs possess a generalisable competence in applying rules, including cases in which a rule's text and purpose point towards different outcomes. Method. Study 1A compared LLM judgments with newly collected human data on published stimuli and matched vignettes created after the models' training cut-offs. Studies 2A/2B compared responses to time-pressure instructions, a manipulation with a mechanistic route to human judgment blocked for LLMs. Study 3 varied reasoning effort, as an analogue for time constrained human judgements. Studies 1B/2B alsovaried system prompt wording and numerical scale anchors. Results LLM judgments closely tracked human judgments for both stimulus sets, while responding in the same unanticipated purposivist direction in the new set as humans did. Sensitivity to text and purpose was robust across prompt variations. Responses to time-pressure instructions were model-specific, suggesting a distinction between conceptual competence and human alignment. Replication of the human pattern was most apparent in models with fewer parameters, and these effects were susceptible to prompt variation. Increasing reasoning effort produced no detectable change in rule application for most models though a significant purposivist trend was observed in higher-effort for GPT-oss and Claude Sonnet 5. Response variance remained lower for LLMs than humans despite our per-model temperature calibration to match human sample variance. Conclusions. Overall, the findings suggest that LLM rule application reflects a generalisable, standing semantic competence that does not typically depend on expanded deliberation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。