arXiv:2608.12306cs.LGcs.AI2026-08中稿 · IJCAI

将稀疏的停止反馈转化为密集成本,提升安全离线强化学习性能

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

论文配图:Redistribution-based Cost Inference Improves Sparse Safe Offline RL
图 1 · 摘自论文原文
  • 通过回报分解将轨迹级停止信号转为每步成本
  • 在高速驾驶与机器人操作任务中违规率显著降低
  • 对数据异质性和标签噪声具有强鲁棒性

安全离线强化学习通常依赖每步的成本标注,但实际中监督者仅提供轨迹级停止反馈:在首次不安全转移时给出二值信号,无具体步骤归因。本文将其建模为时间信用分配问题,提出基于重分配的成本推断(RCI)框架,通过回报分解将稀疏停止反馈转化为密集每步成本,并在此增强数据集上训练约束型离线策略。理论证明回报等价重分配保持了马尔可夫决策过程中的可行策略集与最优拉格朗日值,实现理论无损;实践中则改善了成本评论器的学习条件。在高速公路驾驶与机器人操作任务上的实验表明,相比稀疏与分类器基线,违规率大幅下降,且对异构数据分布和标签噪声具有鲁棒性。

原文摘要 · Abstract (English)

Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset. We show that return-equivalent redistribution preserves the feasible policy set and the optimal Lagrangian in a CMDP, establishing that the transformation is lossless in theory while yielding better-conditioned cost critic learning in practice. Experiments on highway driving and robotic manipulation demonstrate substantially lower violation rates than sparse and classifier-based baselines, with robustness to heterogeneous dataset compositions and label noise.

离线强化学习安全控制成本推断稀疏反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。