用因果分析揭示大模型自我反思背后的可解释行为层级
ReBeCA: Unveiling Interpretable Behavior Hierarchy behind the Iterative Self-Reflection of Language Models with Causal Analysis
- 将自我反思过程建模为因果图,通过三阶段不变因果预测识别真实影响因素
- 发现仅少数语义行为对反思效果有因果影响,且存在层级结构
- 揭示看似正面的行为叠加反而会降低反思效果,适合研究模型可解释性者
尽管自我反思能提升语言模型可靠性,其内在机制仍不清晰,现有分析多基于相关性,难以泛化。为此,我们提出 exttt{ReBeCA}(通过因果分析揭示自我反思行为),将自我反思轨迹建模为因果图,采用三阶段不变因果预测(ICP)流程,识别性能的真实决定因素。关键发现:(1) 模型语义行为以层级方式影响最终反思结果;(2) 反思效果的泛化仅限于少数语义行为;(3) 即便在直接因果因素中,多个看似积极行为的共现也可能削弱反思效能。ICP验证显示,稀疏因果父节点可实现最高49.6%的结构似然提升,且在多种任务中保持稳定,而相关性模式失效。在新数据集上的干预实验确认因果关系具有分布外稳定性(p = .013, η²ₚ = .071)。ReBeCA为剥离自我反思动态中虚假关联、揭示真实因果机制提供了严谨方法。
原文摘要 · Abstract (English)
While self-reflection can enhance language model reliability, its underlying mechanisms remain opaque, with existing analyses often yielding correlation-based insights that fail to generalize. To address this, we introduce \textbf{\texttt{ReBeCA}} (self-\textbf{\texttt{Re}}flection \textbf{\texttt{Be}}havior explained through \textbf{\texttt{C}}ausal \textbf{\texttt{A}}nalysis), a framework that unveils the interpretable behavioral hierarchy governing the self-reflection outcome. By modeling self-reflection trajectories as causal graphs, ReBeCA isolates genuine determinants of performance through a three-stage Invariant Causal Prediction (ICP) pipeline. We establish three critical findings: (1) \textbf{Behavioral hierarchy:} Semantic behaviors of the model influence final self-reflection results hierarchically: directly or indirectly; (2) \textbf{Causation matters:} Generalizability in self-reflection effects is limited to just a few semantic behaviors; (3) \textbf{More $\mathbf{\neq}$ better:} The confluence of seemingly positive semantic behaviors, even among direct causal factors, can impair the efficacy of self-reflection. ICP-based verification identifies sparse causal parents achieving up to $49.6\%$ structural likelihood gains, stable across tasks where correlation-based patterns fail. Intervention studies on novel datasets confirm these causal relationships hold out-of-distribution ($p = .013, η^2_\mathrm{p} = .071$). ReBeCA thus provides a rigorous methodology for disentangling genuine causal mechanisms from spurious associations in self-reflection dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。