让AI理解复杂指令并分步操作3D物体,实现智能交互
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
- 用多模态大模型解析用户指令,分步生成可操作区域
- 构建18万条指令-点云数据集,支持序列化操作推理
- 适合机器人操控、智能助手等需要复杂任务理解的场景
3D可操作性分割旨在将人类指令与3D物体的可触区域关联,用于具身操作。现有方法通常局限于单对象、单可操作性范式,每个可操作类型或显式指令严格对应特定区域,难以处理长时序任务。该范式无法主动推理复杂用户意图,而这些意图常隐含多个顺序可操作性。本文提出序列化3D可操作性推理任务,通过从复杂用户意图中推理并分解为一系列分割图,扩展传统范式。为此,我们构建首个基于指令的可操作性分割基准,涵盖单个和序列化可操作性,包含18万条指令-点云对。基于此基准,我们提出SeqAfford模型,使3D多模态大语言模型具备额外的可操作性分割能力,确保结合世界知识与细粒度可操作性定位的统一框架。我们进一步引入多粒度语言-点云融合模块,实现3D密集预测。大量实验评估表明,该模型优于现有主流方法,并展现出开放世界泛化能力与序列推理能力。
原文摘要 · Abstract (English)
3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance region and are unable to handle long-horizon tasks. Such a paradigm cannot actively reason about complex user intentions that often imply sequential affordances. In this paper, we introduce the Sequential 3D Affordance Reasoning task, which extends the traditional paradigm by reasoning from cumbersome user intentions and then decomposing them into a series of segmentation maps. Toward this, we construct the first instruction-based affordance segmentation benchmark that includes reasoning over both single and sequential affordances, comprising 180K instruction-point cloud pairs. Based on the benchmark, we propose our model, SeqAfford, to unlock the 3D multi-modal large language model with additional affordance segmentation abilities, which ensures reasoning with world knowledge and fine-grained affordance grounding in a cohesive framework. We further introduce a multi-granular language-point integration module to endow 3D dense prediction. Extensive experimental evaluations show that our model excels over well-established methods and exhibits open-world generalization with sequential reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。