arXiv:2604.01345cs.LG2026-04

用马利万微积分高效估计反事实梯度,实现自适应逆强化学习。

Malliavin Calculus for Counterfactual Gradient Estimation in Adaptive Inverse Reinforcement Learning

论文配图:Malliavin Calculus for Counterfactual Gradient Estimation in Adaptive Inverse Reinforcement Learning
图 1 · 摘自论文原文
  • 引入马利万微积分解决反事实梯度估计难题。
  • 在被动观测下实现与标准速率相当的估计精度。
  • 适合研究逆强化学习与强化学习优化的学者。

逆强化学习(IRL)通过观察前向学习者的响应来恢复其损失函数。自适应IRL旨在仅被动观测学习者执行强化学习时的梯度,从而重构其损失函数。本文提出一种基于朗之万过程的新型被动算法,实现自适应IRL。其核心挑战在于:被动算法所需的梯度为反事实梯度——即条件于概率为零事件的梯度。因此,朴素蒙特卡洛估计器效率极低,而常用的核平滑方法则收敛缓慢。本文通过引入马利万微积分,高效估计所需反事实梯度。将反事实条件转化为涉及马利万量的无条件期望比值,从而恢复标准估计速率。推导了通用朗之万结构下的必要马利万导数及其伴随斯科罗霍德积分形式,并给出了可实际执行的算法框架。

原文摘要 · Abstract (English)

Inverse reinforcement learning (IRL) recovers the loss function of a forward learner from its observed responses. Adaptive IRL aims to reconstruct the loss function of a forward learner by passively observing its gradients as it performs reinforcement learning (RL). This paper proposes a novel passive Langevin-based algorithm that achieves adaptive IRL. The key difficulty in adaptive IRL is that the required gradients in the passive algorithm are counterfactual, that is, they are conditioned on events of probability zero under the forward learner's trajectory. Therefore, naive Monte Carlo estimators are prohibitively inefficient, and kernel smoothing, though common, suffers from slow convergence. We overcome this by employing Malliavin calculus to efficiently estimate the required counterfactual gradients. We reformulate the counterfactual conditioning as a ratio of unconditioned expectations involving Malliavin quantities, thus recovering standard estimation rates. We derive the necessary Malliavin derivatives and their adjoint Skorohod integral formulations for a general Langevin structure, and provide a concrete algorithmic approach which exploits these for counterfactual gradient estimation.

逆强化学习马利万微积分梯度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。