无需标注即可分解动态城市场景,实现精准实例级重建与编辑。
UnIRe: Unsupervised Instance Decomposition for Dynamic Urban Scene Reconstruction
- 基于4D超点聚类,利用时空相关性无监督分离动态物体。
- 在多个基准数据集上优于现有方法,支持任意动态类别重建。
- 适合自动驾驶与城市规划等需要灵活编辑的现实应用。
动态城市场景的重建与分解对自动驾驶、城市规划和场景编辑至关重要。然而,现有方法在无手动标注的情况下无法实现实例感知的分解,这限制了实例级场景编辑能力。本文提出UnIRe,一种基于3D高斯溅射(3DGS)的方法,仅使用RGB图像和LiDAR点云,将场景分解为静态背景与独立动态实例。核心是引入4D超点——一种在4D空间中聚类多帧LiDAR点的新表示,通过时空相关性实现无监督实例分离。这些4D超点构成解耦的4D初始化,为训练动态3DGS提供时空初始化,无需边界框或物体模板即可处理任意动态类别。此外,我们在2D和3D空间引入平滑性正则化策略,进一步提升时间稳定性。在基准数据集上的实验表明,该方法在解耦动态场景重建方面优于现有方法,并支持准确灵活的实例级编辑,具备实际应用潜力。
原文摘要 · Abstract (English)
Reconstructing and decomposing dynamic urban scenes is crucial for autonomous driving, urban planning, and scene editing. However, existing methods fail to perform instance-aware decomposition without manual annotations, which is crucial for instance-level scene editing.We propose UnIRe, a 3D Gaussian Splatting (3DGS) based approach that decomposes a scene into a static background and individual dynamic instances using only RGB images and LiDAR point clouds. At its core, we introduce 4D superpoints, a novel representation that clusters multi-frame LiDAR points in 4D space, enabling unsupervised instance separation based on spatiotemporal correlations. These 4D superpoints serve as the foundation for our decomposed 4D initialization, i.e., providing spatial and temporal initialization to train a dynamic 3DGS for arbitrary dynamic classes without requiring bounding boxes or object templates.Furthermore, we introduce a smoothness regularization strategy in both 2D and 3D space, further improving the temporal stability.Experiments on benchmark datasets show that our method outperforms existing methods in decomposed dynamic scene reconstruction while enabling accurate and flexible instance-level editing, making it a practical solution for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。