arXiv:2608.26737cs.CVcs.LG2026-08

用扩散模型统一解决室外激光扫描语义补全难题

Generative Semantic Scene Completion

论文配图:Generative Semantic Scene Completion
图 1 · 摘自论文原文
  • 提出三角色统一扩散框架,从数据生成到补全再到优化
  • 单步推理达38.8% mIoU,超越同类方法2.1个百分点
  • 无需重训练或测试时调整,适合实时自动驾驶系统

室外LiDAR语义场景补全(SSC)需从仅覆盖目标体积1%的稀疏扫描中恢复密集语义体素网格,类别不平衡超过7000倍。本文将SSC重构为生成式语义场景补全(GSSC),采用单一离散扩散框架实现三个功能:首先,通过配对稀疏-稠密场景合成(PS³)生成匹配的稀疏观测与稠密语义补全数据,缓解长尾问题,构建PS³-SemanticKITTI数据集;其次,语义引导生成补全(SGSC)基于鸟瞰视图语义图和稀疏3D特征流,使用多项离散扩散从噪声生成完整场景;第三,同一框架可一步优化已有补全结果,即结构化源离散扩散(S²D²)。S²D²在不重新训练或测试时调整的情况下,提升SGSC自身输出及所有外部基线模型性能。在最强基线模型上,单步无测试增强即达38.8% mIoU(SemanticKITTI隐藏测试集),为当前最佳因果、单扫、单样本结果,较前一公开最优分高2.1个百分点。四步修正+八视角测试增强可达39.2%。

原文摘要 · Abstract (English)

Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. First, paired sparse-dense scene synthesis (PS$^3$) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS$^3$-SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic-guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird's-eye-view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow-matching step: structured source discrete diffusion (S$^2$D$^2$). S$^2$D$^2$ improves the mIoU of SGSC's own output and every external SSC base tested, without base retraining or test-time adaptation. On the strongest base, one step without test-time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single-sweep, single-sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight-view test-time augmentation reach 39.2%, outside that restriction.

语义补全扩散模型点云生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。