arXiv:2511.19004cs.CV2025-11被引 4

用自条件引导生成更真实的文本到激光雷达场景。

A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation

  • 引入自条件表示引导,提升生成细节和几何结构。
  • 在T2nuScenes数据集上优于现有方法,生成场景更真实。
  • 支持多种条件生成任务,适合自动驾驶场景构建。

文本到激光雷达生成可为下游任务定制具有丰富结构和多样场景的3D数据。然而,文本-激光雷达配对数据稀缺导致训练先验不足,生成场景过于平滑;低质量文本描述也会降低生成质量和可控性。本文提出一种文本到激光雷达扩散模型T2LDM,引入自条件表示引导(SCRG)。SCRG通过与真实表示对齐,在训练时为去噪网络提供重建细节的软监督,推理时解耦。该机制使T2LDM能从数据分布中感知丰富几何结构,生成更细致的物体。同时,构建了内容可组合的T2nuScenes基准及可控性度量,分析不同文本提示对生成质量与可控性的影响,提供实用提示范式。此外,设计方向位置先验以缓解街道畸变,进一步提升场景保真度。通过冻结去噪网络学习条件编码器,T2LDM还可支持稀疏到稠密、稠密到稀疏、语义到激光雷达等多任务生成。大量实验表明,无论无条件还是条件生成,T2LDM均达到领先性能。

原文摘要 · Abstract (English)

Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient training priors, generating overly smooth 3D scenes. Moreover, low-quality text descriptions may degrade generation quality and controllability. In this paper, we propose a Text-to-LiDAR Diffusion Model for scene generation, named T2LDM, with a Self-Conditioned Representation Guidance (SCRG). Specifically, SCRG, by aligning to the real representations, provides the soft supervision with reconstruction details for the Denoising Network (DN) in training, while decoupled in inference. In this way, T2LDM can perceive rich geometric structures from data distribution, generating detailed objects in scenes. Meanwhile, we construct a content-composable Text-LiDAR benchmark, T2nuScenes, along with a controllability metric. Based on this, we analyze the effects of different text prompts for LiDAR generation quality and controllability, providing practical prompt paradigms and insights. Furthermore, a directional position prior is designed to mitigate street distortion, further improving scene fidelity. Additionally, by learning a conditional encoder via frozen DN, T2LDM can support multiple conditional tasks, including Sparse-to-Dense, Dense-to-Sparse, and Semantic-to-LiDAR generation. Extensive experiments in unconditional and conditional generation demonstrate that T2LDM outperforms existing methods, achieving state-of-the-art scene generation.

文本生成激光雷达扩散模型3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。