arXiv:2410.09747cs.CVcs.AI2024-10被引 9

t-READi让自动驾驶感知更鲁棒高效,能自动适应传感器差异。

t-READi: Transformer-Powered Robust and Efficient Multimodal Inference for Autonomous Driving

  • 只调整敏感参数,保持其他结构不变,实现自适应推理。
  • 在真实数据下提升平均准确率超6%,推理延迟降低近15倍。
  • 适合部署在传感器常出问题的自动驾驶系统中。

由于自动驾驶车辆广泛采用多模态传感器(如摄像头、激光雷达、雷达),融合其输出以实现鲁棒感知变得至关重要。然而,现有融合方法常依赖两个不切实际的假设:一、所有输入数据分布相似;二、所有传感器始终可用。例如,激光雷达分辨率各异,雷达可能失效,这种变异性常导致融合性能显著下降。为此,我们提出t-READi,一种自适应推理系统,可应对多模态传感数据的变异性,从而实现鲁棒且高效的感知。t-READi识别对变化敏感但结构特定的模型参数,并仅调整这些参数,其余部分保持不变。同时,它采用跨模态对比学习来补偿缺失模态带来的损失。两项机制均兼容现有深度多模态融合方法。大量实验表明,相较于现有方法,t-READi在真实数据与模态变化条件下,平均推理准确率提升超过6%,推理延迟降低近15倍,最坏情况下仅增加5%内存开销。

原文摘要 · Abstract (English)

Given the wide adoption of multimodal sensors (e.g., camera, lidar, radar) by autonomous vehicles (AVs), deep analytics to fuse their outputs for a robust perception become imperative. However, existing fusion methods often make two assumptions rarely holding in practice: i) similar data distributions for all inputs and ii) constant availability for all sensors. Because, for example, lidars have various resolutions and failures of radars may occur, such variability often results in significant performance degradation in fusion. To this end, we present tREADi, an adaptive inference system that accommodates the variability of multimodal sensory data and thus enables robust and efficient perception. t-READi identifies variation-sensitive yet structure-specific model parameters; it then adapts only these parameters while keeping the rest intact. t-READi also leverages a cross-modality contrastive learning method to compensate for the loss from missing modalities. Both functions are implemented to maintain compatibility with existing multimodal deep fusion methods. The extensive experiments evidently demonstrate that compared with the status quo approaches, t-READi not only improves the average inference accuracy by more than 6% but also reduces the inference latency by almost 15x with the cost of only 5% extra memory overhead in the worst case under realistic data and modal variations.

自动驾驶多模态融合自适应推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。