arXiv:2506.20373cs.ROcs.AI2025-06被引 1

让机器人看懂多人协作中谁在做什么、用了什么

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

  • 用视觉语言模型+物体动作识别,给每个人和物分配唯一身份
  • 在倒水、传递、分类任务中准确识别角色与互动关系
  • 适合需要精准理解群体互动的机器人协作场景

我们提出CARMA系统,用于人机群体交互中的情境定位。有效协作需基于对当前人员与物体的一致表征,以及对参与者和操作物体事件的时序抽象。这要求对实体实例进行清晰一致的指代,确保机器人能正确识别并追踪参与者、物体及其交互。CARMA通过唯一标识现实世界中的实体,并将其组织为演员-物体-动作的结构化三元组。我们在三个实验中验证:多人与机器人协作完成倒水、传递和分类任务。结果表明,系统可稳定生成准确的演员-动作-物体三元组,为需要时空推理和情境决策的应用提供可靠基础。

原文摘要 · Abstract (English)

We introduce CARMA, a system for situational grounding in human-robot group interactions. Effective collaboration in such group settings requires situational awareness based on a consistent representation of present persons and objects coupled with an episodic abstraction of events regarding actors and manipulated objects. This calls for a clear and consistent assignment of instances, ensuring that robots correctly recognize and track actors, objects, and their interactions over time. To achieve this, CARMA uniquely identifies physical instances of such entities in the real world and organizes them into grounded triplets of actors, objects, and actions. To validate our approach, we conducted three experiments, where multiple humans and a robot interact: collaborative pouring, handovers, and sorting. These scenarios allow the assessment of the system's capabilities as to role distinction, multi-actor awareness, and consistent instance identification. Our experiments demonstrate that the system can reliably generate accurate actor-action-object triplets, providing a structured and robust foundation for applications requiring spatiotemporal reasoning and situated decision-making in collaborative settings.

人机协作情境理解多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。