构建可理解人与环境互动的3D场景图,让服务机器人更懂社交
Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
- 扩展3D场景图,加入人类属性、动作及远近关系
- 在合成环境中实现对人类活动和关系的准确预测
- 适合研究社交机器人交互与场景理解的学者
理解人与周围环境及他人的互动,是使机器人以符合社会规范且具上下文意识的方式行动的关键。尽管3D场景图已成为场景理解的强大语义表示,但现有方法大多忽略场景中的人类,且因缺乏人类-环境关系标注而受限。此外,现有方法通常仅从单帧图像中捕捉开放词汇关系,难以建模超出观察内容的长距离互动。本文提出社交3D场景图,一种增强型3D场景图表示,通过开放词汇框架捕捉人类及其属性、活动与环境中的局部和远程关系。同时,我们构建了一个新基准,包含带全面人类-场景关系标注的合成环境,以及多样化的查询类型,用于评估3D场景中的社交理解能力。实验表明,该表示显著提升了人类活动预测与人-环境关系推理性能,为社交智能机器人铺平道路。
原文摘要 · Abstract (English)
Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for scene understanding, existing approaches largely ignore humans in the scene, also due to the lack of annotated human-environment relationships. Moreover, existing methods typically capture only open-vocabulary relations from single image frames, which limits their ability to model long-range interactions beyond the observed content. We introduce Social 3D Scene Graphs, an augmented 3D Scene Graph representation that captures humans, their attributes, activities and relationships in the environment, both local and remote, using an open-vocabulary framework. Furthermore, we introduce a new benchmark consisting of synthetic environments with comprehensive human-scene relationship annotations and diverse types of queries for evaluating social scene understanding in 3D. The experiments demonstrate that our representation improves human activity prediction and reasoning about human-environment relations, paving the way toward socially intelligent robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。