arXiv:2606.03915cs.CV2026-06

用局部补丁扩散生成大场景点云,精度与泛化性俱佳。

PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion

论文配图:PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion
图 1 · 摘自论文原文
  • 分块扩散生成局部3D几何,避免全局表示瓶颈
  • 在SemanticKITTI上超越现有方法,50米场景无需重训
  • 适合自动驾驶中长距离点云补全任务

我们提出PatchScene,一种基于扩散模型的大规模激光雷达场景补全新框架。不同于依赖全局隐变量或稠密体素网格的方法,PatchScene采用分块体素扩散范式,在局部3D区域显式生成细粒度几何结构。为保证时空一致性,引入置信度引导的时空融合机制,统一生成重叠补丁与相邻帧。此外,设计环形流扩散策略,利用激光雷达扫描的径向密度模式,从近距到远距逐步传播高保真信息,实现无空间边界场景补全。在SemanticKITTI基准上的大量实验表明,PatchScene在所有标准指标上均达当前最优,不仅几何精度更高,且时间一致性更强。值得注意的是,仅在20米范围内训练的模型可有效泛化至50米场景而无需重新训练,展现出强大的可扩展性与实际应用潜力。

原文摘要 · Abstract (English)

We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent representations or dense voxel grids, PatchScene adopts a patch-based voxel diffusion paradigm that explicitly generates fine-grained geometry within localized 3D regions. To ensure coherent reconstruction at both spatial and temporal scales, we introduce a confidence-guided spatio-temporal fusion mechanism that integrates overlapping patches and adjacent frames in a unified generative process. Furthermore, we design an Annular-Flow diffusion strategy that leverages the radial density pattern of LiDAR scans to progressively propagate high-fidelity information from near-range to far-range regions, enabling spatially unbounded scene completion. Extensive experiments on the SemanticKITTI benchmark demonstrate that PatchScene achieves state-of-the-art performance across all standard metrics, surpassing previous approaches in both geometric accuracy and temporal consistency. Remarkably, the model trained on 20 m LiDAR ranges generalizes effectively to 50 m scenes without retraining, highlighting its strong scalability and generalization capability for real-world autonomous driving applications.

点云补全扩散模型自动驾驶大场景生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。