发现大模型自我反思的关键是动作路由,而非复杂提示或术语。
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
- 通过六组对照实验分离反思的四个组件,精准定位核心机制。
- 动作路由使预测F1提升0.101,而术语和结构化提问无效。
- 适合开发具备自我认知能力的决策类AI系统的研究者。
自反思被普遍认为能提升大模型推理能力,但其关键驱动因素仍不明确。本文通过六条件受控消融实验,分离了证据暴露、诊断支架、分类术语和动作路由四个组件。结果表明:结构化诊断问题相比自由反思无显著增益(F1=0.296 vs 0.297,p=1.000);呈现完整不确定性分类体系但仅保留单一通用动作也未带来提升(ΔF1=+0.008,95% CI重叠)。而类型化动作路由持续带来方向性增益(F1=0.379 vs 0.296),保守估计控制术语影响后增益ΔF1=+0.075,整体相较单次基线显著提升ΔF1=+0.101(95% CI [0.020, 0.185])。该结论在GPT-4o上复现:术语无显著价值(p=0.773),动作路由有显著增益(p=0.025)。增益集中在结构新颖冲突:缅甸(F1: 0.000→0.353)、乌克兰(0.167→0.500),术语条件未能超越通用反思,而动作路由打破退化先验。研究确认类型化动作路由是元认知大模型预测代理的核心设计原则。
原文摘要 · Abstract (English)
Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection ($\text{F1} = 0.296$ vs $0.297$, $p = 1.000$, 95\% CI $[-0.041, +0.040]$). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value ($Δ\text{F1} = +0.008$, overlapping 95\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains ($\text{F1} = 0.379$ vs $0.296$); the conservative estimate controlling for taxonomy vocabulary is $Δ\text{F1} = +0.075$, and the overall gain over the single-shot baseline is significant by bootstrap CI ($Δ\text{F1} = +0.101$, 95\% CI $[+0.020, +0.185]$). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection ($p = 0.773$), while action routing provides significant gains ($p = 0.025$), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar ($\text{F1}: 0.000 \rightarrow 0.353$) and Ukraine ($0.167 \rightarrow 0.500$), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。