T2LDM++用自条件引导生成逼真点云,解决文本-激光雷达数据少的问题。
T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation

- 引入自条件表示引导,通过重建监督让模型学几何特征
- 构建超10万样本的高质量文本-点云数据集,支持多种条件生成
- 可生成带丰富几何细节的点云场景,适合自动驾驶研究
文本到点云生成因缺乏大规模文本-点云配对数据,常导致场景过平滑、可控性差。本文提出T2LDM++,采用自条件表示引导(SCRG)机制,通过引导网络(GN)为去噪网络(DN)提供基于重建的软监督,使DN学习到更精准的几何感知表征,提升DDPM中去噪准确性。该设计轻量高效,推理时解耦,无额外开销。同时,我们基于泛化几何标注策略构建了两个高质量文本-点云基准数据集(>10万样本),并设计可控性度量。此外,引入定向位置先验缓解街道畸变,提升场景保真度。T2LDM++还支持多条件生成(如语义/边界框/鸟瞰图/相机→点云、稀疏→稠密、稠密→稀疏),通过冻结去噪网络学习控制编码器。结合有效先验建模与高质量数据集,模型可在无条件与有条件设置下生成具有丰富几何细节的逼真点云场景。
原文摘要 · Abstract (English)
Recent progress in Text-to-Image generation benefits from large-scale Text-Image pairs. However, the scarcity of Text-LiDAR pairs often causes over-smoothed scenes and limited controllability. In this paper, we rethink the limitations of Text-LiDAR generation task, focusing on alleviating insufficient training priors and constructing controllable Text-LiDAR data. We propose a \textbf{T}ext-\textbf{to}-\textbf{L}iDAR \textbf{D}iffusion \textbf{M}odel for LiDAR scene generation, T2LDM++, with a Self-Conditioned Representation Guidance (SCRG). Specifically, to alleviate object over-smoothing, SCRG employs a Guidance Network (GN) to provide reconstruction-based soft supervision to the Denoising Network (DN). This enables DN to learn geometry-aware representations through reconstruction guidance, leading to more accurate denoising in DDPMs. Meanwhile, through analysis and design, SCRG exhibits more effective and lightweight, while decoupled in inference, avoiding computational overhead. Furthermore, we construct two high-quality Text-LiDAR benchmarks ($>$100K samples) using a generalized strategy of geometric annotations, along with a controllability metric. Moreover, a directional position prior is designed to mitigate street distortion, further improving scene fidelity. Additionally, T2LDM++ supports multiple conditions, including (Semantic, Box, BEV, Camera)-to-LiDAR, Sparse-to-Dense, and Dense-to-Sparse generation, by learning a control encoder via frozen DN. With effective prior modeling and high-quality Text-LiDAR benchmarks, T2LDM++ can generate realistic LiDAR scenes with rich geometric details in unconditional and conditional settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。