arXiv:2509.07463cs.ROcs.AI2025-09

用激光雷达生成逼真图像,让视觉模型在黑暗中也能看清路。

DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving

  • 用条件GAN+精修器将稀疏点云转为稠密彩色图像
  • 在低光照下比纯摄像头方案理解能力提升显著
  • 无需改动现有模型,适合快速集成到自动驾驶系统

确保视觉输入退化时智能车辆与机器人的可靠运行仍是关键挑战。本文提出DepthVision多模态框架,使视觉-语言模型(VLMs)无需架构修改或重新训练即可利用激光雷达数据。该方法通过带有集成精修器的条件生成对抗网络(GAN),将稀疏激光雷达点云合成密集、类RGB图像,并通过标准视觉接口输入现成的VLMs。Luminance-Aware Modality Adaptation(LAMA)模块根据环境光照动态加权合成图像与真实相机图像,有效补偿昏暗或运动模糊等退化。此设计使激光雷达在摄像头失效时可作为即插即用的视觉替代源,显著扩展了现有VLMs的可用范围。我们在真实与模拟数据集上对多个VLMs及安全关键任务进行了评估,包括车端闭环实验。结果表明,在低光照场景下,深度视觉模型的场景理解能力远超仅依赖RGB的基线,同时保持与冻结的VLM架构完全兼容。这些发现证明,基于激光雷达引导的图像合成是将距离感知融入现代视觉-语言系统的实用路径。

原文摘要 · Abstract (English)

Ensuring reliable autonomous operation when visual input is degraded remains a key challenge in intelligent vehicles and robotics. We present DepthVision, a multimodal framework that enables Vision--Language Models (VLMs) to exploit LiDAR data without any architectural changes or retraining. DepthVision synthesizes dense, RGB-like images from sparse LiDAR point clouds using a conditional GAN with an integrated refiner, and feeds these into off-the-shelf VLMs through their standard visual interface. A Luminance-Aware Modality Adaptation (LAMA) module fuses synthesized and real camera images by dynamically weighting each modality based on ambient lighting, compensating for degradation such as darkness or motion blur. This design turns LiDAR into a drop-in visual surrogate when RGB becomes unreliable, effectively extending the operational envelope of existing VLMs. We evaluate DepthVision on real and simulated datasets across multiple VLMs and safety-critical tasks, including vehicle-in-the-loop experiments. The results show substantial improvements in low-light scene understanding over RGB-only baselines while preserving full compatibility with frozen VLM architectures. These findings demonstrate that LiDAR-guided RGB synthesis is a practical pathway for integrating range sensing into modern vision-language systems for autonomous driving.

自动驾驶多模态生成模型激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。