提出新方法,在稀疏数据下仍能准确比较奖励函数差异。
Reward Distance Comparisons Under Transition Sparsity
- 设计可适应稀疏采样的奖励距离度量方法
- 在低覆盖率条件下仍保持高精度比较结果
- 适合数据采集受限的强化学习应用场景
奖励比较对评估不同奖励函数引发的智能体行为差异至关重要。传统方法需先学习最优策略再进行比较,计算成本高且存在安全风险。直接奖励比较虽省去策略学习,但在过渡稀疏场景下表现不佳——因数据收集困难,仅能采样少量转移状态。现有先进方法依赖高覆盖过渡采样,当该条件不满足时,采样与期望转移分布不匹配,导致显著误差。本文提出抗稀疏奖励距离(SRRD)伪度量,无需高过渡覆盖率,可适配多样本分布,常见于过渡稀疏情形。我们提供理论依据说明其鲁棒性,并通过多领域实验验证其实际有效性。
原文摘要 · Abstract (English)
Reward comparisons are vital for evaluating differences in agent behaviors induced by a set of reward functions. Most conventional techniques utilize the input reward functions to learn optimized policies, which are then used to compare agent behaviors. However, learning these policies can be computationally expensive and can also raise safety concerns. Direct reward comparison techniques obviate policy learning but suffer from transition sparsity, where only a small subset of transitions are sampled due to data collection challenges and feasibility constraints. Existing state-of-the-art direct reward comparison methods are ill-suited for these sparse conditions since they require high transition coverage, where the majority of transitions from a given coverage distribution are sampled. When this requirement is not satisfied, a distribution mismatch between sampled and expected transitions can occur, leading to significant errors. This paper introduces the Sparsity Resilient Reward Distance (SRRD) pseudometric, designed to eliminate the need for high transition coverage by accommodating diverse sample distributions, which are common under transition sparsity. We provide theoretical justification for SRRD's robustness and conduct experiments to demonstrate its practical efficacy across multiple domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。