arXiv:2602.02451cs.LGcs.AI2026-02

用偏好学习让实验策略自动进化,比传统方法快70%以上

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization

  • 通过对比干预选择而非奖励数值,学习连续实验决策策略
  • 在多种场景下,相同实验次数下比基线提升70%-71%(p<0.001)
  • 能自主发现关键因果机制,适合需要高效实验设计的领域

发现因果关系依赖受控实验,但实验者面临序列决策问题:每次干预提供信息,应指导下一步行动。传统方法如随机采样、贪婪信息最大化和轮换覆盖均孤立处理每一步决策,无法从经验中学习适应性策略。我们提出主动因果实验家(ACE),将实验设计建模为序列策略。核心洞见是:尽管绝对信息增益随知识积累而递减(导致基于价值的强化学习不稳定),但候选干预间的相对比较始终有意义。ACE通过直接偏好优化,从成对干预比较中学习,而非非平稳的奖励值。在合成基准、物理模拟和经济数据上,ACE在同等干预预算下相比基线实现70-71%的性能提升(p < 0.001,Cohen's d ~ 2)。值得注意的是,学习到的策略自主发现:碰撞器机制需集中干预父变量,这一理论支持的策略完全由经验生成。这表明偏好学习可恢复严谨的实验策略,以学习实现领域自适应,补充理论。

原文摘要 · Abstract (English)

Discovering causal relationships requires controlled experiments, but experimentalists face a sequential decision problem: each intervention reveals information that should inform what to try next. Traditional approaches such as random sampling, greedy information maximization, and round-robin coverage treat each decision in isolation, unable to learn adaptive strategies from experience. We propose Active Causal Experimentalist (ACE), which learns experimental design as a sequential policy. Our key insight is that while absolute information gains diminish as knowledge accumulates (making value-based RL unstable), relative comparisons between candidate interventions remain meaningful throughout. ACE exploits this via Direct Preference Optimization, learning from pairwise intervention comparisons rather than non-stationary reward magnitudes. Across synthetic benchmarks, physics simulations, and economic data, ACE achieves 70-71% improvement over baselines at equal intervention budgets (p < 0.001, Cohen's d ~ 2). Notably, the learned policy autonomously discovers that collider mechanisms require concentrated interventions on parent variables, a theoretically-grounded strategy that emerges purely from experience. This suggests preference-based learning can recover principled experimental strategies, complementing theory with learned domain adaptation.

因果推断实验设计偏好学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。