跨模态3D场景理解新框架,支持缺失数据下的精准检索与定位。
CrossOver: 3D Scene Cross-Modal Alignment
- 通过统一嵌入空间实现多模态柔性对齐,无需对象级语义标注
- 在ScanNet和3RScan上表现优于现有方法,支持缺失模态仍保持鲁棒性
- 适合真实场景中数据不完整时的3D理解任务,如智能导航、数字孪生
多模态3D物体理解受到广泛关注,但现有方法通常假设所有模态数据完整且严格对齐。我们提出CrossOver,一种基于灵活场景级模态对齐的新型跨模态3D场景理解框架。不同于传统方法要求每个物体实例在所有模态下都有对齐数据,CrossOver通过学习统一的、模态无关的嵌入空间,对RGB图像、点云、CAD模型、平面图和文本描述等多模态数据进行松弛约束下的对齐,且无需显式对象语义信息。借助特定维度编码器、多阶段训练流程及涌现的跨模态行为,CrossOver实现了鲁棒的场景检索与物体定位能力,即使部分模态缺失也表现良好。在ScanNet和3RScan数据集上的评估显示其在多种指标上均表现优异,突显了其在真实世界3D场景理解应用中的适应性。
原文摘要 · Abstract (English)
Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment. Unlike traditional methods that require aligned modality data for every object instance, CrossOver learns a unified, modality-agnostic embedding space for scenes by aligning modalities -- RGB images, point clouds, CAD models, floorplans, and text descriptions -- with relaxed constraints and without explicit object semantics. Leveraging dimensionality-specific encoders, a multi-stage training pipeline, and emergent cross-modal behaviors, CrossOver supports robust scene retrieval and object localization, even with missing modalities. Evaluations on ScanNet and 3RScan datasets show its superior performance across diverse metrics, highlighting the adaptability for real-world applications in 3D scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。