arXiv:2603.23950cs.RO2026-03被引 1

机器人通过观察人与物体交互自动判断何时主动协助,更像真人搭档。

Event-Driven Proactive Assistive Manipulation with Grounded Vision-Language Planning

  • 基于环境状态变化触发主动协助,而非等待指令
  • 通过前后状态快照识别任务进展,提升正确干预率23%
  • 适合需要自然协作的现实场景,如家庭助手机器人

协作操作中的辅助通常由用户指令触发,依赖高阶推理。但在流畅的人类团队合作中,伙伴常根据已发生动作的结果推断下一步帮助行为,而非等待指令。受此启发,我们提出从请求驱动转向事件驱动的主动辅助模式:机器人通过检测人-物交互引发的工作区状态转移来触发行动,而非依赖用户任务指令。为此,我们设计了一套事件驱动框架,利用事件监控器追踪交互进度,在事件完成后提取稳定的前后状态快照以表征状态变迁。基于这些快照,规划器分析隐含的状态变化,推断任务级目标并决定是否干预;若需介入,则生成一系列助行动作。为确保输出可执行且可验证,所有动作均限定在一组动作原语和通过整数ID引用的对象范围内。我们在真实桌面数字积木协作任务上评估该框架,结果表明,显式前后状态变化证据能提升可解场景下的主动完成率,并在不可解场景中实现恰当等待。

原文摘要 · Abstract (English)

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we introduce a shift from request-driven assistance to event-driven proactive assistance, where robot actions are initiated by workspace state transitions induced by human--object interactions rather than user-provided task instructions. To this end, we propose an event-driven framework that tracks interaction progress with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. Given the stabilized snapshots, the planner analyzes the implied state transition to infer a task-level goal and decide whether to intervene; if so, it generates a sequence of assistive actions. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs. We evaluate the framework on a real tabletop number-block collaboration task, demonstrating that explicit pre/post state-change evidence improves proactive completion on solvable scenes and appropriate waiting on unsolvable ones.

主动协助视觉语言规划协作机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。