让机器人在陌生环境中用双手协同完成复杂操作,不需重新训练。
Scene-agnostic Hierarchical Bimanual Task Planning via Visual Affordance Reasoning
- 通过视觉分析生成可交互的3D位置点,定位物体操作位。
- 设计双臂协作子目标规划器,实现空间合理、动作紧凑的双臂指令。
- 适合需要双手协同的智能机器人任务,如装配或搬移复杂物品。
在开放环境中,具身智能体需将高层指令转化为可执行的物理行为,常需双手协同操作。尽管现有基础模型具备强大的语义推理能力,但传统机器人任务规划仍以单手为主,难以应对场景无关设置下的空间、几何与协调挑战。本文提出统一框架,实现无需场景依赖的双臂任务规划,将高层语义推理与三维物理行为执行相衔接。核心包含三个模块:视觉点定位(VPG)从单张图像中检测相关物体并生成世界对齐的交互点;双臂子目标规划(BSP)基于空间邻近性与跨物体可达性,生成紧凑且运动中立的子目标,挖掘协同双臂操作机会;交互点驱动双臂提示(IPBP)将子目标绑定至结构化技能库,生成满足手状态与功能约束的同步单/双臂动作序列。该框架使智能体可在杂乱、未见过的场景中规划出语义合理、物理可行且可并行执行的双臂行为。实验表明,该方法能生成连贯、可行且简洁的双臂计划,并在不重新训练的前提下泛化至复杂场景,验证了其在双臂任务中鲁棒的场景无关功能推理能力。
原文摘要 · Abstract (English)
Embodied agents operating in open environments must translate high-level instructions into grounded, executable behaviors, often requiring coordinated use of both hands. While recent foundation models offer strong semantic reasoning, existing robotic task planners remain predominantly unimanual and fail to address the spatial, geometric, and coordination challenges inherent to bimanual manipulation in scene-agnostic settings. We present a unified framework for scene-agnostic bimanual task planning that bridges high-level reasoning with 3D-grounded two-handed execution. Our approach integrates three key modules. Visual Point Grounding (VPG) analyzes a single scene image to detect relevant objects and generate world-aligned interaction points. Bimanual Subgoal Planner (BSP) reasons over spatial adjacency and cross-object accessibility to produce compact, motion-neutralized subgoals that exploit opportunities for coordinated two-handed actions. Interaction-Point-Driven Bimanual Prompting (IPBP) binds these subgoals to a structured skill library, instantiating synchronized unimanual or bimanual action sequences that satisfy hand-state and affordance constraints. Together, these modules enable agents to plan semantically meaningful, physically feasible, and parallelizable two-handed behaviors in cluttered, previously unseen scenes. Experiments show that it produces coherent, feasible, and compact two-handed plans, and generalizes to cluttered scenes without retraining, demonstrating robust scene-agnostic affordance reasoning for bimanual tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。