arXiv:2409.10095cs.CV2024-09被引 1

统一编码器融合多种视觉任务,提升自动驾驶决策能力。

Human Insights Driven Latent Space for Different Driving Perspectives: A Unified Encoder for Efficient Multi-Task Inference

  • 设计统一编码器,联合训练深度、位姿、3D流等多任务感知。
  • 在转向预测任务中,冻结编码器性能优于微调版和ImageNet预训练模型。
  • 适合需要高效多任务推理的自动驾驶系统研发者。

自动驾驶系统需全面理解环境,依赖于感知、规划与控制所需的视觉特征提取。然而,仅基于单任务目标或通用数据集训练的模型常缺乏复杂驾驶场景下的上下文信息。本文提出一个统一编码器,联合训练城市驾驶关键任务:深度估计、位姿估计、3D场景流估计,以及语义、实例、全景和运动分割。通过整合多样视觉线索(类比人类感知机制),编码器捕捉丰富特征,增强导航相关预测。在转向估计任务上评估其性能,利用密集潜在空间表示。为实现高效多任务学习,引入多尺度特征网络用于位姿估计,并采用多主干教师模型的知识蒸馏。实验表明:(1) 统一编码器在所有感知任务上表现优异,具备强泛化能力;(2) 冻结的统一编码器在转向预测中优于其微调版本及ImageNet预训练模型。结果凸显任务特定视觉特征的重要性,验证了多任务学习在推进自动驾驶系统中的潜力。更多细节与预训练模型见https://hi-computervision.github.io/uni-encoder/。

原文摘要 · Abstract (English)

Autonomous driving systems require a comprehensive understanding of the environment, achieved by extracting visual features essential for perception, planning, and control. However, models trained solely on single-task objectives or generic datasets often lack the contextual information needed for robust performance in complex driving scenarios. In this work, we propose a unified encoder trained on multiple computer vision tasks crucial for urban driving, including depth, pose, and 3D scene flow estimation, as well as semantic, instance, panoptic, and motion segmentation. By integrating these diverse visual cues-similar to human perceptual mechanisms-the encoder captures rich features that enhance navigation-related predictions. We evaluate the model on steering estimation as a downstream task, leveraging its dense latent space. To ensure efficient multi-task learning, we introduce a multi-scale feature network for pose estimation and apply knowledge distillation from a multi-backbone teacher model. Our findings highlight two key findings: (1) the unified encoder achieves competitive performance across all visual perception tasks, demonstrating strong generalization capabilities; and (2) for steering estimation, the frozen unified encoder-leveraging dense latent representations-outperforms both its fine-tuned counterpart and the same frozen model pretrained on generic datasets like ImageNet. These results underline the significance of task-specific visual features and demonstrate the promise of multi-task learning in advancing autonomous driving systems. More details and the pretrained model are available at https://hi-computervision.github.io/uni-encoder/.

自动驾驶多任务学习视觉编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。