arXiv:2606.21292cs.CV2026-06被引 1

用概率方法将2D模型嵌入转为稳定3D语义表示,适合处理噪声和模糊场景。

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models

论文配图:Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models
图 1 · 摘自论文原文
  • 通过变分推理建模多视角观测,融合相对姿态推断3D语义状态。
  • 在新视角上预测语义特征,相比简单拼接提升3D语义稳定性。
  • 不依赖特定主干网络,支持语言对齐与自监督嵌入,适合开放词汇3D理解。

我们提出Casper3D,一种轻量级概率框架,可将带有噪声的多视角2D基础模型嵌入转化为潜在的3D语义表示。将视图级语义特征视为底层3D语义状态的噪声观测,利用结合相对姿态的集合变分模型进行多视角推理。Casper3D通过从新视角预测被遮蔽的语义观测进行训练,同时保持与视觉和文本语义空间的一致性,实现开放词汇的3D理解。该框架与主干网络无关,适用于语言对齐和自监督嵌入。实验表明,在模糊和噪声环境下,其生成的3D语义比简单多视角池化更稳定。

原文摘要 · Abstract (English)

We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-level semantic features as noisy observations of an underlying 3D semantic state and infer this state with a set-based variational model that incorporates relative pose during multi-view reasoning. Casper3D is trained by predicting held-out semantic observations from novel viewpoints, while remaining aligned with visual and text semantic spaces for open-vocabulary 3D understanding. The framework is backbone-agnostic and applies to both language-aligned and self-supervised embeddings. Experiments show that Casper3D produces more stable 3D semantics than simple multi-view pooling, especially in ambiguous and noisy settings.

3D生成概率建模多视角推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。