无需标注数据,用渲染生成视角实现鸟瞰图语义分割
RendBEV: Semantic Novel View Synthesis for Self-Supervised Bird's Eye View Segmentation

- 用可微体素渲染从2D分割图生成语义视角提供自监督信号
- 零样本下即达到可比性能,小标注量时显著提升精度
- 适合缺乏标注数据的自动驾驶场景,尤其适合预训练
鸟瞰图(BEV)语义地图近年来成为辅助与自动驾驶任务的重要环境表征。然而现有方法多依赖大规模标注数据的全监督训练。本文提出RendBEV,一种基于可微体素渲染的自监督训练方法,利用2D语义分割模型生成的语义视角作为监督信号。该方法实现零样本下的BEV语义分割,已具备竞争力;在少量标注条件下进行微调时性能显著提升,且在全部标签可用时达到新最优水平。
原文摘要 · Abstract (English)
Bird's Eye View (BEV) semantic maps have recently garnered a lot of attention as a useful representation of the environment to tackle assisted and autonomous driving tasks. However, most of the existing work focuses on the fully supervised setting, training networks on large annotated datasets. In this work, we present RendBEV, a new method for the self-supervised training of BEV semantic segmentation networks, leveraging differentiable volumetric rendering to receive supervision from semantic perspective views computed by a 2D semantic segmentation model. Our method enables zero-shot BEV semantic segmentation, and already delivers competitive results in this challenging setting. When used as pretraining to then fine-tune on labeled BEV ground-truth, our method significantly boosts performance in low-annotation regimes, and sets a new state of the art when fine-tuning on all available labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。