提出新方法评估机器人抓取策略,更准反映任务进展。
Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation

- 用任务完成度构建贝尔曼算子,应对有限步长偏差
- 在模拟抓取与折叠任务中,显著降低截断偏差
- 适合评估带恢复行为的稀疏奖励策略
策略评估是机器人策略开发与部署的核心环节。在现代操作任务中,奖励稀疏、评估轨迹的任务进展常非单调(因策略具恢复行为),且轨迹长度有限,导致标准方法依赖的无限时域假设失效。本文提出基于存活性(liveness)的离线策略评估框架,将评估视为任务完成问题,构建保守型固定点价值函数,有效抵御有限时域截断带来的偏差。理论分析显示该算子具有收缩性,可编码任务进展并缓解截断偏差。我们在两个模拟操作任务上使用视觉-语言-动作模型与扩散策略,在布料折叠任务上使用人类示范进行验证。实验表明,该方法更准确反映任务进度,显著优于经典基线如TD(0)与蒙特卡洛评估。
原文摘要 · Abstract (English)
Policy evaluation is a fundamental component of the development and deployment pipeline for robotic policies. In modern manipulation systems, this problem is particularly challenging: rewards are often sparse, task progression of evaluation rollouts are often non-monotonic as the policies exhibit recovery behaviors, and evaluation rollouts are necessarily of finite length. This finite length introduces truncation bias, breaking the infinite-horizon assumptions underlying standard methods relying on Bellman equations/principle of optimality. In this work, we propose a framework for offline policy evaluation from sparse rewards based on a liveness-based Bellman operator. Our formulation interprets policy evaluation as a task-completion problem and yields a conservative fixed-point value function that is robust to finite-horizon truncation. We analyze the theoretical properties of the proposed operator, including contraction guarantees, and show how it encodes task progression while mitigating truncation bias. We evaluate our method on two simulated manipulation tasks using both a Vision-Language-Action model and a diffusion policy, and a cloth folding task using human demonstrations. Empirical results demonstrate that our approach more accurately reflects task progress and substantially reduces truncation bias, outperforming classical baselines such as TD(0) and Monte Carlo policy evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。