构建3D场景中人类行为理解新任务,提出首个专用基础模型。
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
- 将3D场景与人体运动融合输入大语言模型,建模人-场景交互。
- 在HIS-Bench上超越现有模型,显著提升问答准确率。
- 适合研究具身智能、世界模型与多模态理解的学者使用。
我们提出一项新任务——人-场景问答(HIS-QA),用于评估具身智能体对3D场景中人类状态与行为的理解能力。给定人体运动序列,该任务要求智能体理解人类行为、推理环境关系,并回答相关问题。为此,我们构建了HIS-Bench,一个涵盖感知到常识推理与规划的多模态基准。对多种视觉-语言模型在该基准上的评估显示其在处理此类任务时存在明显局限。为此,我们提出HIS-GPT,首个面向人-场景理解的基础模型。它将3D场景上下文与人体运动动态融入大语言模型,并引入专门机制捕捉人-场景交互。大量实验表明,HIS-GPT在HIS-QA任务上达到新最佳性能。本工作旨在推动3D场景中人类行为分析的研究,促进具身智能与世界模型的发展。代码与数据:https://github.com/ZJHTerry18/HumanInScene。
原文摘要 · Abstract (English)
We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-related questions within the scene. To support this new task, we present HIS-Bench, a multimodal benchmark that systematically evaluates HIS understanding across a broad spectrum, from basic perception to commonsense reasoning and planning. Our evaluation of various vision-language models on HIS-Bench reveals significant limitations in their ability to handle HIS-QA tasks. To this end, we propose HIS-GPT, the first foundation model for HIS understanding. HIS-GPT integrates 3D scene context and human motion dynamics into large language models while incorporating specialized mechanisms to capture human-scene interactions. Extensive experiments demonstrate that HIS-GPT sets a new state-of-the-art on HIS-QA tasks. We hope this work inspires future research on human behavior analysis in 3D scenes, advancing embodied AI and world models. The codes and data: https://github.com/ZJHTerry18/HumanInScene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。