用连续概率场统一建模动态场景中物体的可见性与身份,提升跨时空一致性。
Consistent Instance Field for Dynamic Scene Understanding
- 以可变形3D高斯为基元,联合编码辐射与语义信息
- 在HyperNeRF和Neu3D上实现新视角全景分割与开放词汇4D查询领先性能
- 适合需要长期一致物体识别的自动驾驶与视频理解任务
我们提出一致实例场(Consistent Instance Field),一种用于动态场景理解的连续且概率化的时空表征。与依赖离散追踪或视图相关特征的现有方法不同,本方法通过为每个时空点建模占据概率和条件实例分布,将可见性与持久物体身份解耦。为此,我们引入基于可变形3D高斯的新实例嵌入表示,联合编码辐射与语义信息,并通过可微投影直接从输入的RGB图像和实例掩码中学习。此外,我们设计了新的机制以校准每个高斯体的身份并将其重采样至语义活跃区域,确保空间时间上的一致实例表征。在HyperNeRF和Neu3D数据集上的实验表明,该方法在新视角全景分割和开放词汇4D查询任务上显著优于现有最先进方法。
原文摘要 · Abstract (English)
We introduce Consistent Instance Field, a continuous and probabilistic spatio-temporal representation for dynamic scene understanding. Unlike prior methods that rely on discrete tracking or view-dependent features, our approach disentangles visibility from persistent object identity by modeling each space-time point with an occupancy probability and a conditional instance distribution. To realize this, we introduce a novel instance-embedded representation based on deformable 3D Gaussians, which jointly encode radiance and semantic information and are learned directly from input RGB images and instance masks through differentiable rasterization. Furthermore, we introduce new mechanisms to calibrate per-Gaussian identities and resample Gaussians toward semantically active regions, ensuring consistent instance representations across space and time. Experiments on HyperNeRF and Neu3D datasets demonstrate that our method significantly outperforms state-of-the-art methods on novel-view panoptic segmentation and open-vocabulary 4D querying tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。