用相机+陀螺仪融合提升道路分类鲁棒性,适应复杂环境变化。
A New Dataset and Framework for Robust Road Surface Classification via Camera-IMU Fusion
- 通过双向交叉注意力融合图像与惯性数据,动态调整模态权重。
- 在新数据集ROAD上比现有方法提升11.6个百分点,少数类识别更准。
- 适合需要低成本、高适应性的道路监测系统开发者使用。
道路表面分类(RSC)是环境感知预测性维护系统的关键技术。然而,现有方法常因传感模态和数据集多样性不足,在狭窄工况下表现不佳。本文提出一种多模态框架,融合图像与惯性测量数据,采用轻量级双向交叉注意力模块,并引入自适应门控层以应对领域偏移时的模态贡献调整。针对现有基准数据集缺乏变异性的问题,我们构建了新数据集ROAD,包含三个互补子集:(i) 实际场景中通过工业级数据记录器同步的RGB-IMU流,覆盖多样光照、天气与路面条件;(ii) 大规模纯视觉子集,用于评估恶劣光照与异构采集设备下的鲁棒性;(iii) 合成子集,用于研究难以实测的分布外场景。实验表明,该方法在PVS基准上相比前序最优模型提升1.4个百分点,在多模态ROAD子集上提升11.6个百分点,且少数类F1分数始终更高。框架在夜间、暴雨及混合路面过渡等挑战性视觉条件下也保持稳定性能。结果表明,结合低成本相机与陀螺仪传感器,辅以多模态注意力机制,可为道路理解提供可扩展、鲁棒的解决方案,尤其适用于环境差异大且成本受限的地区。
原文摘要 · Abstract (English)
Road surface classification (RSC) is a key enabler for environment-aware predictive maintenance systems. However, existing RSC techniques often fail to generalize beyond narrow operational conditions due to limited sensing modalities and datasets that lack environmental diversity. This work addresses these limitations by introducing a multimodal framework that fuses images and inertial measurements using a lightweight bidirectional cross-attention module followed by an adaptive gating layer that adjusts modality contributions under domain shifts. Given the limitations of current benchmarks, especially regarding lack of variability, we introduce ROAD, a new dataset composed of three complementary subsets: (i) real-world multimodal recordings with RGB-IMU streams synchronized using a gold-standard industry datalogger, captured across diverse lighting, weather, and surface conditions; (ii) a large vision-only subset designed to assess robustness under adverse illumination and heterogeneous capture setups; and (iii) a synthetic subset generated to study out-of-distribution generalization in scenarios difficult to obtain in practice. Experiments show that our method achieves a +1.4 pp improvement over the previous state-of-the-art on the PVS benchmark and an +11.6 pp improvement on our multimodal ROAD subset, with consistently higher F1-scores on minority classes. The framework also demonstrates stable performance across challenging visual conditions, including nighttime, heavy rain, and mixed-surface transitions. These findings indicate that combining affordable camera and IMU sensors with multimodal attention mechanisms provides a scalable, robust foundation for road surface understanding, particularly relevant for regions where environmental variability and cost constraints limit the adoption of high-end sensing suites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。