arXiv:2601.22206cs.LGstat.ME2026-01

在观测噪声和分布偏移下,用因果建模提升离线模仿学习的鲁棒性。

Causal Imitation Learning Under Measurement Error and Distribution Shift

  • 将噪声观测视为代理变量,构建具有因果解释的策略目标
  • 在无奖励、无交互情况下,可从示范数据中恢复目标策略
  • 适用于医疗时序数据等存在测量误差的场景,适合关注可靠性研究者

我们研究在部分决策相关状态仅通过噪声观测且训练与部署分布可能变化的离线模仿学习(IL)场景。此类设置会引入虚假的状态-动作关联,导致标准行为克隆(BC)——无论是否基于原始观测或忽略它们——在分布偏移下收敛于系统性偏差的策略。我们提出一个受因果关系建模启发的通用框架,生成具有因果解释的目标,并对分布偏移具有鲁棒性。基于近端因果推断思想,我们提出 exttt{CausIL},将噪声状态观测视为代理变量,并给出了在无奖励、无交互专家查询条件下目标策略可识别的条件。我们为离散与连续状态空间设计了估计器;在连续情形下,使用核希尔伯特空间(RKHS)上的对抗过程学习所需参数。我们在来自 PhysioNet/Computing in Cardiology Challenge 2019 的半模拟纵向数据上评估 exttt{CausIL},结果表明其相比 BC 基线具有更强的分布偏移鲁棒性。

原文摘要 · Abstract (English)

We study offline imitation learning (IL) when part of the decision-relevant state is observed only through noisy measurements and the distribution may change between training and deployment. Such settings induce spurious state-action correlations, so standard behavioral cloning (BC) -- whether conditioning on raw measurements or ignoring them -- can converge to systematically biased policies under distribution shift. We propose a general framework for IL under measurement error, inspired by explicitly modeling the causal relationships among the variables, yielding a target that retains a causal interpretation and is robust to distribution shift. Building on ideas from proximal causal inference, we introduce \texttt{CausIL}, which treats noisy state observations as proxy variables, and we provide identification conditions under which the target policy is recoverable from demonstrations without rewards or interactive expert queries. We develop estimators for both discrete and continuous state spaces; for continuous settings, we use an adversarial procedure over RKHS function classes to learn the required parameters. We evaluate \texttt{CausIL} on semi-simulated longitudinal data from the PhysioNet/Computing in Cardiology Challenge 2019 cohort and demonstrate improved robustness to distribution shift compared to BC baselines.

模仿学习因果推断分布偏移医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。