新方法同时学习奖励分布与风险感知策略,更真实还原专家行为。
Distributional Inverse Reinforcement Learning
- 用分布框架联合建模奖励不确定性和回报分布
- 在合成数据和MuJoCo任务上实现顶尖性能
- 适合需要理解行为风险与复杂决策的场景
我们提出一种离线逆强化学习的分布式框架,联合建模奖励函数的不确定性与回报的完整分布。不同于传统方法仅恢复确定性奖励或匹配期望回报,本方法通过最小化一阶随机占优违规,引入扭曲风险度量(DRMs)到策略学习中,从而捕捉专家行为的更丰富结构,实现奖励分布与分布感知策略的联合恢复。该框架适用于行为分析与风险敏感型模仿学习。理论分析表明算法收敛速度为$/mathcal{O}(\varepsilon^{-2})$。在合成基准、真实神经行为数据及MuJoCo控制任务上的实验表明,该方法能恢复表达性强的奖励表示,并达到当前最优性能。
原文摘要 · Abstract (English)
We propose a distributional framework for offline Inverse Reinforcement Learning (IRL) that jointly models uncertainty over reward functions and full distributions of returns. Unlike conventional IRL approaches that recover a deterministic reward estimate or match only expected returns, our method captures richer structure in expert behavior, particularly in learning the reward distribution, by minimizing first-order stochastic dominance (FSD) violations and thus integrating distortion risk measures (DRMs) into policy learning, enabling the recovery of both reward distributions and distribution-aware policies. This formulation is well-suited for behavior analysis and risk-aware imitation learning. Theoretical analysis shows that the algorithm converges with $\mathcal{O}(\varepsilon^{-2})$ iteration complexity. Empirical results on synthetic benchmarks, real-world neurobehavioral data, and MuJoCo control tasks demonstrate that our method recovers expressive reward representations and achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。