让大模型学会在实验中重构假设空间,突破原有认知框架。
Separable Pathways for Causal Reasoning: How Architectural Scaffolding Enables Hypothesis-Space Restructuring in LLM Agents

- 用结构化图与动态行为分离设计,分别管理推理路径和假设更新。
- 在1085次实验中,新架构使准确率提升94%,关键在于及时发现旧假设失效。
- 适合研究可解释性、因果推理的AI系统开发者或认知科学交叉研究者。
通过实验与干预进行因果发现是稳健问题求解的基础,这不仅需要在固定框架内更新信念,更需重构假设空间本身——而当前AI代理缺乏这种能力,当证据要求其使用未曾构建过的表征时便束手无策。本文将发展科学中的blicket检测范式拓展至AI代理,引入针对假设空间重构的架构支撑。所提组合式架构包含两个独立模块:上下文图(context graphs)将探索过程建模为带类型的状态机,动态行为(dynamic behaviors)则实时监测证据,识别当前假设空间不足时主动扩展。在1,085次实验测试中,二者贡献正交:上下文图主导切换后假设空间内的推理质量,贡献了94%的准确率提升;动态行为则决定推理可行性,通过检测状态变迁防止过早承诺于过时假设。
原文摘要 · Abstract (English)
Causal discovery through experimentation and intervention is fundamental to robust problem solving. It requires not just updating beliefs within a fixed framework but revising the hypothesis space itself, a capacity current AI agents lack when evidence demands representations they have not previously constructed. We extend the blicket detector paradigm from developmental science to test this capacity in AI agents equipped with architectural scaffolding that targets hypothesis-space restructuring. Our compositional architecture has two discrete components: context graphs, which structure exploration as typed state machines, and dynamic behaviors, which monitor for evidence that the current hypothesis space is inadequate and expand it at runtime. Across 1,085 experimental trials, these components make orthogonal contributions: context graphs drive reasoning quality within the post-switch hypothesis space, accounting for 94\% of the accuracy gain, while dynamic behaviors drive reasoning eligibility by detecting regime changes and preventing premature commitment to outdated hypotheses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。