提出3D多对象推理分割任务,提升复杂场景理解能力。
Multimodal 3D Reasoning Segmentation with Complex Scenes
- 设计多对象查询的3D推理网络MORE3D,捕捉空间关系
- 在复杂多物场景中实现高精度分割与文本解释生成
- 构建大规模基准ReasonSeg3D,适合3D视觉语言研究者
多模态学习的发展推动了3D场景理解在具身智能等实际任务中的进步。然而,现有方法普遍存在两个问题:缺乏对人类意图的交互与推理能力;且多聚焦于单类物体与简化文本描述,忽视多物体间复杂的空间关系。为此,本文提出3D推理分割任务,支持在多物体复杂场景中生成3D分割掩码与包含空间关系的详细文本解释。我们构建了大规模高质量基准ReasonSeg3D,集成3D分割掩码、3D空间关系及自动生成的问答对。同时设计新颖的3D推理网络MORE3D,可处理多对象查询,学习空间关系细节并用于捕获物体空间信息与推理文本输出。大量实验表明,MORE3D在复杂多物体3D场景的推理与分割上表现优异。所建ReasonSeg3D为未来3D推理分割研究提供宝贵平台。数据与代码将公开。
原文摘要 · Abstract (English)
The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of reasoning ability for interaction and interpretation of human intentions and 2) they focus on scenarios with single-category objects and over-simplified textual descriptions and neglect multi-object scenarios with complicated spatial relations among objects. We address the above challenges by proposing a 3D reasoning segmentation task for reasoning segmentation with multiple objects in scenes. The task allows producing 3D segmentation masks and detailed textual explanations as enriched by 3D spatial relations among objects. To this end, we create ReasonSeg3D, a large-scale and high-quality benchmark that integrates 3D segmentation masks and 3D spatial relations with generated question-answer pairs. In addition, we design MORE3D, a novel 3D reasoning network that works with queries of multiple objects and is tailored for 3D scene understanding. MORE3D learns detailed explanations on 3D relations and employs them to capture spatial information of objects and reason textual outputs. Extensive experiments show that MORE3D excels in reasoning and segmenting complex multi-object 3D scenes. In addition, the created ReasonSeg3D offers a valuable platform for future exploration of 3D reasoning segmentation. The data and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。