让轮椅机械臂与移动机器人临时牵手,完成双手协作任务。
JOIN: Anchor-Grasp-Conditioned Joining via Opposition, Inference, and Navigation for Bimanual Assistive Manipulation

- 用视觉语言模型+几何工具分三步规划双手配合动作
- 在20次测试中成功19次,纠错次数显著减少
- 适合残障人士的双臂辅助操作,尤其适合空间受限场景
辅助移动与操作平台正日益成为帮助残障人士恢复独立性的关键。尽管许多日常活动已可通过单臂系统完成,但诸如开罐、倒液体、端托盘和基本备餐等任务本质上需双手协同,单臂系统难以应对。为避免在轮椅上加装第二臂带来的能耗增加、成本上升及空间占用问题,本文提出一种异构、按需启用的双臂系统:由轮椅固定的锚定臂在需要时与召唤来的移动操作臂(互补臂)临时连接。核心挑战是‘双臂对接’——锚定臂已固定抓握,互补臂需判断自身位置与抓取目标以完成任务。本文将该问题分解为规划、驱动、抓取三阶段,并证明结合视觉语言模型(VLM)与标准几何工具可提供足够任务级知识,解决典型双臂日常生活任务。所提系统JOIN引入(i)基于轮椅参考系的对立评分,以及(ii)任务条件下的方向性操作能力。在Kinova Gen3锚定臂与Hello Robot Stretch~3互补臂上评估,针对同物体与异物体任务,JOIN完成尝试19/20次,优于现有方法的14/20,且显著降低操作员修正需求。
原文摘要 · Abstract (English)
Assistive mobility and manipulation platforms have received increasing attention as a means of restoring independence to individuals with disabilities. While effective for many basic activities of daily living (ADLs), a significant percentage of everyday tasks such as opening a jar, pouring a liquid, lifting a tray, or basic meal preparation, is fundamentally bimanual and remains out of reach for any single-arm system. Adding a second arm to a wheelchair is impractical, due to the additional power draw, cost, and the loss of space required for transfers and mobility. We instead propose a heterogeneous, on-demand bimanual system, in which a wheelchair-mounted anchor arm is joined when needed by a summoned mobile manipulator that serves as a complement arm. The central technical problem, which we call bimanual joining, is conditional: the anchor has already committed to a grasp, and the complement arm must choose where to stand and what to grasp to complete the task. We formulate bimanual joining as a three-phase decomposition (plan, drive, grasp) and show that a vision-language model (VLM), coupled with standard geometric tools, provides task-level knowledge sufficient to solve a representative class of bimanual ADLs. Our system JOIN, contributes (i) a wheelchair-referenced opposition score, and (ii) task-conditioned directional manipulability. We evaluate JOIN on a Kinova Gen3 anchor and a Hello Robot Stretch~3 complement on representative same-object and different-object tasks. JOIN accomplished more attempts (19/20) than state-of-the-art methods (14/20) and required markedly less correction by the operator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。