测试大模型在空间经济推理中的结构化干预效果,发现不同架构模型对同一方法反应相反。
Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

- 用五种推理干预对比两个模型在八类问题上的表现
- 架构差异导致干预效果反转:承诺提升标准模型但损害优化模型
- 推理模型更易受对抗攻击,且难执行已知策略
我们研究结构化推理干预是否能提升大语言模型的战略经济推理能力,以及其效果是否依赖模型架构。以霍特林线性城市模型为诊断工具,评估 GPT-4.1-mini(标准指令跟随模型)和 GPT-5-mini(推理优化模型)在五种条件下——无引导基线及四种推理干预——的表现,涵盖八道题目,涉及演绎与归纳推理,三种提示框架,每条件重复三次,共获得720个独立评分。结果显示,引导类型与模型架构存在显著交叉作用(t(7)=4.79, p=0.002, d=1.69):承诺引导提升标准模型(+0.21),却使推理模型下降(-0.63);而原则性分离则相反(-0.40 对 +0.31)。两种交叉均显著(承诺:p=0.040;分离:p=0.002),且在全部八题中方向一致性达7/8。对抗性压力测试损害两模型,推理模型受损更严重(-1.47 对 -0.57;p=0.038),且损伤与基础难度负相关(R²=0.36, p=0.014)。此外,两模型均存在声明-执行差距,识别正确策略的比率远高于实际执行率;原则性分离完全弥合了推理模型的此差距,但无干预能改善标准模型。
原文摘要 · Abstract (English)
We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture. Using Hotelling's linear city model as a diagnostic vehicle, we evaluate GPT-4.1-mini (a standard instruction-following model) and GPT-5-mini (a reasoning-optimized model) under five conditions - an unscaffolded baseline and four reasoning interventions - across eight questions spanning deductive and abductive reasoning, three prompt framings, and three repetitions per condition, yielding 720 individually judged responses. We find a statistically significant crossover interaction between scaffolding type and model architecture ($t(7) = 4.79$, $p = 0.002$, $d = 1.69$): commitment scaffolding improves the standard model ($+0.21$) while degrading the reasoning model ($-0.63$), and principled separation shows the opposite pattern ($-0.40$ vs. $+0.31$). Both crossovers are individually significant (commitment: $p = 0.040$; separation: $p = 0.002$) and hold across all eight questions with 7/8 directional consistency. Adversarial stress-testing harms both models, with $2.6\times$ greater degradation for the reasoning model ($-1.47$ vs. $-0.57$; $p = 0.038$), and the damage correlates negatively with baseline difficulty ($R^2 = 0.36$, $p = 0.014$). We further document a persistent declarative-procedural gap in which both models identify correct strategies at rates far exceeding their ability to execute them; separation fully closes this gap for the reasoning model while no intervention helps the standard model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。