arXiv:2507.14303cs.CV2025-07

用语义分割提升自动驾驶环境理解能力

Semantic Segmentation based Scene Understanding in Autonomous Vehicles

  • 基于BDD100k数据集,设计高效分割模型
  • 不同主干网络显著影响分割精度
  • 适合关注自动驾驶视觉感知的研究者

近年来,人工智能在解决复杂任务方面展现出巨大潜力,使机器能在关键场景中自主决策,减少对人类专业知识的依赖。深度学习作为主流人工智能技术,在自动驾驶领域应用尤为有效。本文提出多个高效模型,通过语义分割实现自动驾驶场景理解,并基于BDD100k数据集进行验证。研究还对比了多种主干网络(Backbone)作为编码器的效果,结果表明主干网络选择对模型性能有显著影响。更好的语义分割表现提升了对周围环境的理解能力。最终,模型在准确率、平均交并比(mean IoU)和损失函数上均取得改进。

原文摘要 · Abstract (English)

In recent years, the concept of artificial intelligence (AI) has become a prominent keyword because it is promising in solving complex tasks. The need for human expertise in specific areas may no longer be needed because machines have achieved successful results using artificial intelligence and can make the right decisions in critical situations. This process is possible with the help of deep learning (DL), one of the most popular artificial intelligence technologies. One of the areas in which the use of DL is used is in the development of self-driving cars, which is very effective and important. In this work, we propose several efficient models to investigate scene understanding through semantic segmentation. We use the BDD100k dataset to investigate these models. Another contribution of this work is the usage of several Backbones as encoders for models. The obtained results show that choosing the appropriate backbone has a great effect on the performance of the model for semantic segmentation. Better performance in semantic segmentation allows us to understand better the scene and the environment around the agent. In the end, we analyze and evaluate the proposed models in terms of accuracy, mean IoU, and loss function, and the results show that these metrics are improved.

语义分割自动驾驶深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。