arXiv:2604.08726cs.RO2026-04被引 2

用视觉语言模型让双臂机器人理解任务意图,精准规划抓取位置和分工。

Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning

  • 融合多视角图像与3D信息,用VLM识别任务相关的抓取区域和手臂分配。
  • 在9个真实任务中成功率超几何与语义基线,最高提升18%。
  • 适合需要理解任务目标的双臂机器人研究与应用。

双臂操作需同时判断在物体何处交互以及由哪只手臂执行动作,这涉及抓取位姿与手臂分配的联合推理问题,仅依赖几何信息的规划器因缺乏任务意图的语义理解而无法解决。现有方法或仅将抓取预测视为粗略部件分割,或依赖几何启发式进行手臂分配,未能联合建模任务相关的接触区域与手臂分配。本文将双臂操作重构为联合抓取定位与手臂分配问题,提出一种分层框架,利用视觉语言模型(VLM)实现跨物体类别与任务描述的泛化,无需类别特定训练。方法将多视角RGB-D观测融合为一致的3D场景表示,生成全局6-DoF抓取候选,并通过查询VLM获取每个物体上任务相关的抓取区域及对应的手臂分配,从而在保证几何有效性的同时尊重任务语义。我们在双臂平台上评估了九个真实世界操作任务,涵盖平行操作、协同稳定、工具使用和人类交接四类。结果表明,该方法在任务导向抓取上的成功率持续高于几何与语义基线,证明显式的语义推理有助于在非结构化环境中实现可靠的双臂操作。

原文摘要 · Abstract (English)

Bimanual manipulation requires reasoning about where to interact with an object and which arm should perform each action, a joint affordance localization and arm allocation problem that geometry-only planners cannot resolve without semantic understanding of task intent. Existing approaches either treat affordance prediction as coarse part segmentation or rely on geometric heuristics for arm assignment, failing to jointly reason about task-relevant contact regions and arm allocation. We reframe bimanual manipulation as a joint affordance localization and arm allocation problem and propose a hierarchical framework for task-aware bimanual affordance prediction that leverages a Vision-Language Model (VLM) to generalize across object categories and task descriptions without requiring category-specific training. Our approach fuses multi-view RGB-D observations into a consistent 3D scene representation and generates global 6-DoF grasp candidates, which are then spatially and semantically filtered by querying the VLM for task-relevant affordance regions on each object, as well as for arm allocation to the individual objects, thereby ensuring geometric validity while respecting task semantics. We evaluate our method on a dual-arm platform across nine real-world manipulation tasks spanning four categories: parallel manipulation, coordinated stabilization, tool use, and human handover. Our approach achieves consistently higher task success rates than geometric and semantic baselines for task-oriented grasping, demonstrating that explicit semantic reasoning over affordances and arm allocation helps enable reliable bimanual manipulation in unstructured environments.

双臂操作视觉语言模型抓取规划语义推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。