用单张摄像头图像生成交通场景自然语言描述,提升自动驾驶环境理解能力
Vision-Based Natural Language Scene Understanding for Autonomous Driving: An Extended Dataset and a New Model for Traffic Scene Description Generation
- 结合混合注意力机制提取空间与语义特征
- 在新构建数据集上CIDEr达45.2,显著优于基线模型
- 适合自动驾驶场景理解与多模态交互研究者使用
交通场景理解对自动驾驶车辆准确感知和解析环境至关重要,有助于确保安全导航。本文提出一种新框架,将单张前视摄像头图像转换为简洁的自然语言描述,有效捕捉空间布局、语义关系及驾驶相关线索。所提模型采用混合注意力机制增强空间与语义特征提取,并融合特征生成上下文丰富、细节详实的场景描述。为解决该领域专用数据集稀缺问题,基于BDD100K数据集构建了新数据集,并提供详细构建指南。研究还深入讨论相关评估指标,确定适用于该任务的最佳度量方式。通过CIDEr和SPICE等量化指标以及人工评估进行广泛验证,结果表明所提模型在新数据集上表现优异,有效实现预期目标。
原文摘要 · Abstract (English)
Traffic scene understanding is essential for enabling autonomous vehicles to accurately perceive and interpret their environment, thereby ensuring safe navigation. This paper presents a novel framework that transforms a single frontal-view camera image into a concise natural language description, effectively capturing spatial layouts, semantic relationships, and driving-relevant cues. The proposed model leverages a hybrid attention mechanism to enhance spatial and semantic feature extraction and integrates these features to generate contextually rich and detailed scene descriptions. To address the limited availability of specialized datasets in this domain, a new dataset derived from the BDD100K dataset has been developed, with comprehensive guidelines provided for its construction. Furthermore, the study offers an in-depth discussion of relevant evaluation metrics, identifying the most appropriate measures for this task. Extensive quantitative evaluations using metrics such as CIDEr and SPICE, complemented by human judgment assessments, demonstrate that the proposed model achieves strong performance and effectively fulfills its intended objectives on the newly developed dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。