用反事实流优化离线策略,不越界也能提升表现
Counterfactual Transport Flows for Offline Conservative Trajectory Refinement
- 基于潜在空间检索高反馈轨迹,生成弱监督信号
- 在D4RL上提升历史行为表现,抗过拟合能力强
- 适合需要可解释改进路径的离线决策场景
离线强化学习仅通过历史数据中的回报等可观测结果来改进策略,核心难点在于避免在数据分布外过度外推。本文提出反事实传输流框架,以源轨迹为条件进行离线决策的轨迹精炼。给定低反馈候选轨迹,通过在潜在轨迹空间中检索邻近且任务相关反馈更高的轨迹,构建局部偏好对,作为保守精炼的弱监督信号。该框架学习实例特定的精炼方向:推理时,通过调节精炼强度参数控制候选轨迹的移动距离,实现保留原行为与强改进之间的权衡。在D4RL基准测试(含AntMaze和MuJoCo任务)上,方法有效利用历史回报作为世界反馈改进行为,同时提供可解释的轨迹级优化路径。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) offers a path to policy improvement from logged data alone, using historical returns or other measurable outcomes as world feedback. A key difficulty is improving observed behavior without extrapolating beyond what the offline data supports. We propose \emph{counterfactual transport flows}, a source-conditioned trajectory refinement framework for offline decision-making guided by world feedback. Given a low-feedback candidate trajectory, we construct local preference pairs from offline data by retrieving nearby trajectories in latent trajectory space with higher task-specific feedback, and use them as weak supervision for conservative refinement. The framework learns instance-specific refinement directions: at inference time, a refinement strength parameter controls how far the candidate trajectory is transported, enabling a trade-off between preserving the original behavior and applying stronger improvement. Experiments on D4RL benchmarks, including AntMaze and MuJoCo tasks, show that our method improves behavior from historical returns as world feedback, while providing interpretable trajectory-level refinement paths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。