arXiv:2509.15693cs.CVcs.MM2025-09NeurIPS

用结构化场景提升3D点云与文本对齐效果

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

论文配图:SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions
图 1 · 摘自论文原文
  • 构建含空间关系的多物体场景,配以语言模型优化描述
  • 在多个数据集上实现零样本分类显著提升
  • 适合需要空间推理与跨模态对齐的研究者

三维文本对比学习中,整体大于部分之和。我们提出SceneForge,一种通过结构化多物体场景组合增强3D点云与文本对齐的新框架。该方法利用单个3D形状构建具有明确空间关系的多物体场景,并配以大语言模型优化的连贯多对象描述。通过引入这些结构化、组合式样本进行对比训练,有效缓解大规模3D-文本数据集稀缺问题,显著提升数据复杂度与多样性。系统性研究了关键设计因素,包括每场景物体数量、组合样本占比及场景构建策略。大量实验表明,SceneForge在ModelNet、ScanObjNN、Objaverse-LVIS、ScanNet上的零样本分类任务,以及ShapeNetPart上的少样本部件分割任务中均取得显著性能提升。其组合增强方法具备模型无关性,在多种编码器架构上持续增益。此外,还提升了ScanQA上的3D视觉问答性能,对日益复杂的检索场景具有强泛化能力,并能根据文本指令精准调整空间布局,展现空间推理能力。

原文摘要 · Abstract (English)

The whole is greater than the sum of its parts-even in 3D-text contrastive learning. We introduce SceneForge, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured multi-object scene compositions. SceneForge leverages individual 3D shapes to construct multi-object scenes with explicit spatial relations, pairing them with coherent multi-object descriptions refined by a large language model. By augmenting contrastive training with these structured, compositional samples, SceneForge effectively addresses the scarcity of large-scale 3D-text datasets, significantly enriching data complexity and diversity. We systematically investigate critical design elements, such as the optimal number of objects per scene, the proportion of compositional samples in training batches, and scene construction strategies. Extensive experiments demonstrate that SceneForge delivers substantial performance gains across multiple tasks, including zero-shot classification on ModelNet, ScanObjNN, Objaverse-LVIS, and ScanNet, as well as few-shot part segmentation on ShapeNetPart. SceneForge's compositional augmentations are model-agnostic, consistently improving performance across multiple encoder architectures. Moreover, SceneForge improves 3D visual question answering on ScanQA, generalizes robustly to retrieval scenarios with increasing scene complexity, and showcases spatial reasoning capabilities by adapting spatial configurations to align precisely with textual instructions.

3D对齐空间推理多模态生成增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。