arXiv:2502.01896cs.CVcs.RO2025-02被引 5

让激光雷达模型抗噪更强,提升自动驾驶感知可靠性

INTACT: Inducing Noise Tolerance through Adversarial Curriculum Training for LiDAR-based Safety-Critical Perception and Autonomy

  • 用元学习生成关键区域图,指导模型逐步对抗噪声
  • 在KITTI上使追踪准确率提升9.6%,噪声下达12.4%
  • 适合自动驾驶等对安全要求高的3D感知场景

本文提出INTACT,一种两阶段框架,提升深度神经网络在安全关键感知任务中对噪声激光雷达数据的鲁棒性。通过结合元学习与对抗课程训练(ACT),INTACT系统性应对3D点云中的数据污染与稀疏问题。元学习阶段使教师网络获得任务无关先验,生成能识别关键数据区域的鲁棒显著图;对抗课程训练阶段利用这些显著图,逐步向学生网络引入更复杂的噪声模式,实现精准扰动和抗噪能力提升。在KITTI、Argoverse和ModelNet40等多个数据集上的评估显示,INTACT在目标检测、跟踪与分类任务中性能全面提升,最多提升20%。其中,KITTI多目标追踪准确率(MOTA)从64.1%提升至75.1%(+9.6%),高斯噪声下从52.5%升至73.7%(+12.4%);平均精度(mAP)从59.8%增至69.8%(+10%),噪声下从49.3%增至70.9%(+21.6%)。该框架为资源受限的安全关键系统提供了高效可扩展的解决方案。

原文摘要 · Abstract (English)

In this work, we present INTACT, a novel two-phase framework designed to enhance the robustness of deep neural networks (DNNs) against noisy LiDAR data in safety-critical perception tasks. INTACT combines meta-learning with adversarial curriculum training (ACT) to systematically address challenges posed by data corruption and sparsity in 3D point clouds. The meta-learning phase equips a teacher network with task-agnostic priors, enabling it to generate robust saliency maps that identify critical data regions. The ACT phase leverages these saliency maps to progressively expose a student network to increasingly complex noise patterns, ensuring targeted perturbation and improved noise resilience. INTACT's effectiveness is demonstrated through comprehensive evaluations on object detection, tracking, and classification benchmarks using diverse datasets, including KITTI, Argoverse, and ModelNet40. Results indicate that INTACT improves model robustness by up to 20% across all tasks, outperforming standard adversarial and curriculum training methods. This framework not only addresses the limitations of conventional training strategies but also offers a scalable and efficient solution for real-world deployment in resource-constrained safety-critical systems. INTACT's principled integration of meta-learning and adversarial training establishes a new paradigm for noise-tolerant 3D perception in safety-critical applications. INTACT improved KITTI Multiple Object Tracking Accuracy (MOTA) by 9.6% (64.1% -> 75.1%) and by 12.4% under Gaussian noise (52.5% -> 73.7%). Similarly, KITTI mean Average Precision (mAP) rose from 59.8% to 69.8% (50% point drop) and 49.3% to 70.9% (Gaussian noise), highlighting the framework's ability to enhance deep learning model resilience in safety-critical object tracking scenarios.

激光雷达抗噪训练自动驾驶3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。