arXiv:2506.23434cs.CVcs.RO2025-06NeurIPS被引 14

提出高效潜空间流匹配框架,让激光雷达世界模型跨域通用且大幅降低标注数据依赖。

Towards foundational LiDAR world models with efficient latent flow matching

  • 基于潜变量条件流匹配,提升模型压缩与训练效率。
  • 仅用5%标注数据即超越先前模型,跨域迁移性能提升11%绝对值。
  • 适合自动驾驶、机器人等需低标注成本的场景使用。

基于激光雷达的世界模型相比图像模型具有更结构化和几何感知的表征能力。然而,现有激光雷达世界模型训练范围狭窄,仅在特定领域表现良好。我们首次系统性研究了三个高难度跨域迁移场景:(i) 户外到户内泛化,(ii) 稀疏光束与密集光束适应,(iii) 非语义到语义迁移。在不同微调数据量下,实验表明单一预训练模型相比从零训练可实现最高11%的绝对性能提升(相对提升83%),并在36组对比中30次优于从零训练。该动态学习的迁移能力显著降低了对人工标注数据的依赖:本方法仅需之前模型5%的标注数据即可超越现有语义占用预测模型。此外,我们发现当前激光雷达世界模型存在数据压缩不足和训练目标低效的问题。为此,我们提出基于潜变量条件流匹配(CFM)的框架,在仅用一半训练数据的情况下达到最优重建精度,压缩率比先前方法高6倍。模型在未来轨迹条件下的语义占用预测任务上达到最先进性能,计算效率提高23倍(帧率提升28倍);在语义占用预测任务上也达最先进水平,计算效率提高2倍(帧率提升1.1倍)。

原文摘要 · Abstract (English)

LiDAR-based world models offer more structured and geometry-aware representations than their image-based counterparts. However, existing LiDAR world models are narrowly trained; each model excels only in the domain for which it was built. Can we develop LiDAR world models that exhibit strong transferability across multiple domains? We conduct the first systematic domain transfer study across three demanding scenarios: (i) outdoor to indoor generalization, (ii) sparse-beam & dense-beam adaptation, and (iii) non-semantic to semantic transfer. Given different amounts of fine-tuning data, our experiments show that a single pre-trained model can achieve up to 11% absolute improvement (83% relative) over training from scratch and outperforms training from scratch in 30/36 of our comparisons. This transferability of dynamic learning significantly reduces the reliance on manually annotated data for semantic occupancy forecasting: our method exceed the previous semantic occupancy forecasting models with only 5% of the labeled training data required by prior models. We also observed inefficiencies of current LiDAR world models, mainly through their under-compression of LiDAR data and inefficient training objectives. To address this, we propose a latent conditional flow matching (CFM)-based frameworks that achieves state-of-the-art reconstruction accuracy using only half the training data and a compression ratio 6 times higher than that of prior methods. Our model achieves SOTA performance on future-trajectory-conditioned semantic occupancy forecasting while being 23x more computationally efficient (a 28x FPS speedup); and achieves SOTA performance on semantic occupancy forecasting while being 2x more computationally efficient (a 1.1x FPS speedup).

激光雷达世界模型跨域迁移高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。