arXiv:2609.02830cs.RO2026-09

提出三维评估框架,检验激光雷达语义分割在真实场景下的部署能力

Toward Robust LiDAR Semantic Segmentation for Real-World Deployment: Evaluation under Coarse Labels, Adverse Conditions, and Domain Shifts

论文配图:Toward Robust LiDAR Semantic Segmentation for Real-World Deployment: Evaluation under Coarse Labels, Adverse Conditions, and Domain Shifts
图 1 · 摘自论文原文
  • 设计三维度评估:粗粒度标签、八类传感器退化、跨数据集泛化
  • 发现细粒度榜单排名与安全关键性能不一致,模型在恶劣条件下普遍降级
  • 适用于自动驾驶系统研发者,推动从实验性能向实际部署的转变

基于激光雷达的语义分割是自动驾驶与移动机器人核心感知模块。尽管当前先进方法在标准基准上表现优异,但现有评估仍局限于清洁、单一领域设置和精细标签体系,未能全面评估实际部署可行性。真实系统需应对安全关键标签语义、传感条件劣化及跨域差异,而目前尚无统一评估协议涵盖三方面。本文提出结构化评估协议,从三个互补维度检验激光雷达语义分割模型的部署就绪度:(i) 与自动驾驶安全优先级对齐的粗粒度标签评估,揭示标签粒度对不同方法的影响;(ii) 八类激光雷达退化模拟真实大气、几何与传感器劣化;(iii) 无需适应的跨数据集域泛化。评估包含在嵌入式Jetson AGX Orin平台上的推理速度,直接反映部署约束。结果表明,细粒度基准排名未必反映安全相关性能,所有方法在退化条件下均显著降级,且架构依赖性明显,当前域泛化仍不足以支持可靠部署。研究揭示了基准性能与部署就绪之间的具体差距,并提供了更贴近实际的评估参考。

原文摘要 · Abstract (English)

LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despite the strong performance of recent state-of-the-art methods on standard benchmarks, existing evaluation protocols remain focused on clean, single-domain settings and fine-grained label taxonomies, leaving deployment readiness largely unassessed. Real-world systems must handle safety-critical label semantics, degraded sensing conditions, and cross-domain variability, yet no unified protocol currently addresses all three aspects together. In this paper, we propose a structured evaluation protocol that assesses the deployment readiness of LiDAR semantic segmentation models along three complementary dimensions: (i) coarse-label evaluation aligned with autonomous driving safety priorities, revealing how label granularity affects different methods; (ii) robustness under eight types of LiDAR corruptions designed to emulate real-world atmospheric, geometric, and sensor degradations; and (iii) domain generalization across datasets without adaptation. The evaluation includes inference speed measured on an embedded Jetson AGX Orin platform, directly reflecting deployment constraints. Our results show that fine-grained benchmark rankings do not always reflect safety-relevant performance, that all methods experience substantial degradation under corruptions with architecture-dependent robustness characteristics, and that current domain generalization remains insufficient for reliable deployment. These findings expose concrete gaps between benchmark performance and deployment readiness, and provide a reference protocol for more practically grounded evaluation of LiDAR semantic segmentation.

激光雷达语义分割部署评估鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。