自动构建高质量3D对话数据集,解决视角和指代模糊问题。
Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
- 用规则+多模态模型自动生成无歧义的3D对话数据。
- 产出超200万样本,覆盖5类物体中心问答任务。
- 适合研究3D视觉语言模型的开发者使用。
3D多模态大模型仍落后于2D模型,主要因高质量3D场景-对话数据集稀缺。以往方法依赖昂贵的人工标注,且存在视角模糊(空间描述依赖未知相机位姿)与对象指代模糊(非唯一描述混淆目标与干扰项)两大问题。为此,我们提出全自动流水线,将原始3D扫描转换为无歧义、高质量的对话数据,成本仅为之前的几分之一。该流水线融合基于规则的约束与2D MLLMs、LLMs,实现无需人工干预的可控、可扩展生成。包含四个阶段:(1) 收集物体、帧、场景级的元注释;(2) 构建场景图并修正关系以捕捉邻近物体关系;(3) 生成排他性、紧凑的对象指代描述;(4) 多任务生成合成多样化对话。该流程系统性修复源数据缺陷,最终产出Disc3D数据集,涵盖25,000个混合3D场景,超过200万样本,覆盖场景、视角、物体描述、视觉定位及五类物体中心问答任务。大量实验表明,使用Disc3D训练可在多个公开基准和自建的Disc3D-QA任务上获得一致且显著提升。代码、数据与模型将公开。
原文摘要 · Abstract (English)
3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leave two key ambiguities unresolved: viewpoint ambiguity, where spatial language presumes unknown camera poses, and object referring ambiguity, where non-exclusive descriptions blur the line between targets and distractors. We therefore present a fully automated pipeline that converts raw 3D scans into unambiguous, high-quality dialogue data at a fraction of the previous cost. By synergizing rule-based constraints with 2D MLLMs and LLMs, the pipeline enables controllable, scalable generation without human intervention. The pipeline comprises four stages: (1) meta-annotation collection harvesting object-, frame-, and scene-level captions, (2) scene graph construction with relation correction to capture proximal object relations, (3) discriminative object referring that generates exclusive and compact descriptions, and (4) multi-task data generation synthesizing diverse dialogues. Our pipeline systematically mitigates inherent flaws in source datasets and produces the final Disc3D dataset, over 2 million samples in 25K hybrid 3D scenes, spanning scene, view, and object captioning, visual grounding, and five object-centric QA tasks. Extensive experiments demonstrate that training with Disc3D yields consistent, significant improvements on both public benchmarks and our multifaceted Disc3D-QA tasks. Code, data, and models will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。