用可解释的优先级引导强化学习,让机械臂更高效地在杂乱中搜寻物体。
XPG-RL: Reinforcement Learning with Explainable Priority Guidance for Efficiency-Boosted Mechanical Search
- 基于视觉输入动态选择抓取、移障、调视角等动作优先级。
- 长任务效率提升4.5倍,真实场景成功率显著高于基线方法。
- 适合需要高鲁棒性与可解释性的复杂机械操作任务。
在杂乱环境中,自主机械臂执行机械搜寻(MS)仍面临长期规划与遮挡下状态估计的挑战。本文提出XPG-RL,一种基于原始感官输入的可解释优先级引导强化学习框架,使智能体通过任务驱动的动作优先级机制和上下文感知切换策略,高效完成搜寻任务。该策略从一组离散动作原语(如目标抓取、遮挡物移除、视角调整)中动态选择,并优化策略输出自适应阈值以控制动作选择。感知模块融合RGB-D数据与语义、几何特征,生成结构化场景表示用于决策。仿真与真实世界实验表明,XPG-RL在任务成功率与运动效率上均优于基线方法,长任务效率最高提升4.5倍。结果证明,结合领域知识与可学习决策策略,能实现鲁棒高效的机器人操作。
原文摘要 · Abstract (English)
Mechanical search (MS) in cluttered environments remains a significant challenge for autonomous manipulators, requiring long-horizon planning and robust state estimation under occlusions and partial observability. In this work, we introduce XPG-RL, a reinforcement learning framework that enables agents to efficiently perform MS tasks through explainable, priority-guided decision-making based on raw sensory inputs. XPG-RL integrates a task-driven action prioritization mechanism with a learned context-aware switching strategy that dynamically selects from a discrete set of action primitives such as target grasping, occlusion removal, and viewpoint adjustment. Within this strategy, a policy is optimized to output adaptive threshold values that govern the discrete selection among action primitives. The perception module fuses RGB-D inputs with semantic and geometric features to produce a structured scene representation for downstream decision-making. Extensive experiments in both simulation and real-world settings demonstrate that XPG-RL consistently outperforms baseline methods in task success rates and motion efficiency, achieving up to 4.5$\times$ higher efficiency in long-horizon tasks. These results underscore the benefits of integrating domain knowledge with learnable decision-making policies for robust and efficient robotic manipulation. The project page for XPG-RL is https://yitingzhang1997.github.io/xpgrl/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。