通过2D-3D联合自监督学习,让模型自动学会空间感知能力。
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- 用2D与3D数据联合训练,实现跨模态自监督学习。
- 在3D场景理解上比顶尖单模模型提升14.2%和4.8%。
- 适合做空间认知、点云理解及开放世界感知的研究者。
人类通过多感官协同学习抽象概念,一旦形成,可从单一模态回忆。受此启发,我们提出Concerto,一种极简的空间认知模拟方法,结合3D内模态自蒸馏与2D-3D跨模态联合嵌入。尽管结构简单,Concerto仍能学习出更连贯、信息量更高的空间特征,零样本可视化表现优异。在线性探测中,其性能分别优于当前最优的2D与3D自监督模型14.2%和4.8%,也超过二者特征拼接。全微调下,Concerto在多个场景理解基准上刷新纪录(如ScanNet上达80.7% mIoU)。我们还提出了针对视频提升点云的变体,并设计线性投影器将表示映射至CLIP语言空间,实现开放世界感知。结果表明,Concerto能生成具有精细几何与语义一致性的空间表征。
原文摘要 · Abstract (English)
Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2% and 4.8%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP's language space, enabling open-world perception. These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。