通过接触进度引导,提升机器人抓取任务的执行成功率与效率
HCPG-Flow:Hierarchical Contact-Progress Guidance for Flow-Policy Robot Manipulation

- 引入分层接触进度引导,在接触后切换至任务进展评估
- 在十项仿真任务中成功率达9.5%提升,物理实验减少17.4%完成步数
- 适合需要高精度操作的机器人抓取与复杂动作规划场景
流策略(Flow policy)可表示机器人操作中的多模态动作分布,但每个控制步骤只能执行一个动作。当采样多个动作候选时,基于评判器的排序依赖于候选动作的价值估计,而这些估计在回放缓冲区中可能表现不足。本文提出HCPG-Flow,一种在推理阶段使用的解析选择器,增强SAC-Flow的性能,同时保持其演员与评判器目标不变。该方法在接触后从末端执行器接近转为任务进展评估,根据任务相关距离的一阶下降量对每个动作提案进行评分,对候选集内得分标准化,并执行温度控制的动作嵌入。在十个仿真任务中,HCPG-Flow在两个基准测试上均提升平均成功率,其中在Maniskill上取得9.5个百分点的增益;四个真实物理任务中也表现出高成功率,且成功完成步数减少17.4%。
原文摘要 · Abstract (English)
Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. When several proposals are sampled, critic-based ranking makes data collection depend on value estimates over candidate actions that may be weakly represented in replay. We introduce HCPG-Flow, an analytic rollout-time selector that augments SAC-Flow with hierarchical, object-centric contact-progress guidance while preserving its actor and critic objectives. HCPG switches from end-effector approach to task progress after contact, scores each proposal by the first-order reduction of a task-relevant distance, standardizes scores within the candidate set, and executes a temperature-controlled action embedding. Across ten simulated tasks, HCPG improves mean success over SAC-Flow on both benchmarks, including a 9.5 percentage-point gain on Maniskill. Four physical tasks further show high success with a 17.4% reduction in successful completion steps.Project page: https://hitxraz.github.io/HCPG-Flow/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。