用轨迹增强提升离线强化学习性能,尤其适合数据少且质量差的场景。
Trajectory-Level Data Augmentation for Offline Reinforcement Learning

- 基于任务结构和奖励与价值函数关系设计轨迹增强方法
- 在有限低质轨迹下训练出更优的离线策略,性能显著提升
- 适用于低维、部分可观测等复杂现实场景,理论与实证兼备
我们提出一种面向离线强化学习的数据增强方法,受主动定位问题启发。该方法使模型能在少量次优轨迹上训练出有效的离策略策略。通过利用任务结构、奖励与值函数间的几何关系及记录策略的数学特性,设计了基于轨迹的增强技术。数据收集阶段支持次优策略的记录,提升了数据质量,从而改善离线强化学习性能。我们在不同维度的定位任务及部分可观测条件下,提供了理论依据并进行了充分的实证验证。
原文摘要 · Abstract (English)
We propose a data augmentation method for offline reinforcement learning, motivated by active positioning problems. Particularly, our approach enables the training of off-policy models from a limited number of suboptimal trajectories. We introduce a trajectory-based augmentation technique that exploits task structure and the geometric relationship between rewards, value functions, and mathematical properties of logging policies. During data collection, our augmentation supports suboptimal logging policies, leading to higher data quality and improved offline reinforcement learning performance. We provide theoretical justification for these strategies and validate them empirically across positioning tasks of varying dimensionality and under partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。