让3D大模型学会物体间的精细比较,提升几何理解能力。
Beyond Single Object: Learning 3D Relations with Large Language Models

- 构建多物体3D对比数据集MO3D,引导模型进行细粒度物体间推理。
- 提出Multi-3DLLM模型,在保持局部几何结构的同时建模物体间关系。
- 在形状匹配与变化描述任务中表现优异,适合需要空间推理的应用场景。
现有3D-LLM主要聚焦单个物体或场景描述,难以处理物体间的详细对比。本文提出一个新框架,包含三个部分:(1) MO3D(多物体3D)指令数据集,要求进行细粒度多物体比较;(2) Multi-3DLLM,采用最小化补丁交互变换器(PIT),在保留局部几何信息的同时建模物体间与物体内部关系;(3) 两个应用导向基准测试(形状配对、变化描述),用于检验几何理解能力。近期3D-LLM和2D-VLM在这些任务上表现不佳,缺乏以比较为核心的设计与几何感知。相比之下,基于混合数据训练的Multi-3DLLM学习到几何推理能力,在MO3D上超越所有基线,并实现向单物体分类任务的正向迁移。
原文摘要 · Abstract (English)
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for detailed object-level reasoning across multiple objects with three components: (1) MO3D (Multi-Object in 3D), an instruction dataset requiring fine-grained multi-object comparison; (2) Multi-3DLLM, using a minimal Patch-Interaction Transformer (PIT) that models inter-/intra-object relationships while preserving local geometry; (3) Mini-apps, two application-driven benchmarks (Shape Mating, Change Captioning) that probe geometric understanding for practical use. Recent 3D-LLMs and 2D-VLMs perform poorly on these tasks, lacking both comparison-centric design and geometric awareness. In contrast, Multi-3DLLM trained on our mixture data learns geometric reasoning, surpasses all baselines on MO3D, and provides positive transfer to single-object classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。