提出首个融合多维度引导的行车物体重要性估计模型,显著提升自动驾驶决策能力。
On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down Guidance
- 融合驾驶意图、语义上下文和交通规则三重自上而下引导,结合底层特征
- 在新构建的TOI数据集上达到23.1%的平均精度提升,超越当前最优模型
- 适合自动驾驶感知与决策系统研究者参考,尤其关注场景理解与注意力机制
本文针对驾驶视角视频中行车物体重要性估计问题展开研究。该任务对实现更安全、更智能的驾驶系统至关重要,但当前公开的大规模数据集稀缺,且现有方法多仅依赖自下而上的特征或单一方向引导,难以应对复杂多变的交通场景。为此,本文构建了名为Traffic Object Importance(TOI)的新大规模数据集,并提出首个融合多层自上而下引导(包括驾驶意图、语义上下文和交通规则)与自下而上特征的模型。实验表明,该模型在标准评测中相较近期最优方法(Goal)实现了23.1%的平均精度(AP)提升,验证了多维度引导的有效性。
原文摘要 · Abstract (English)
This paper addresses the problem of on-road object importance estimation, which utilizes video sequences captured from the driver's perspective as the input. Although this problem is significant for safer and smarter driving systems, the exploration of this problem remains limited. On one hand, publicly-available large-scale datasets are scarce in the community. To address this dilemma, this paper contributes a new large-scale dataset named Traffic Object Importance (TOI). On the other hand, existing methods often only consider either bottom-up feature or single-fold guidance, leading to limitations in handling highly dynamic and diverse traffic scenarios. Different from existing methods, this paper proposes a model that integrates multi-fold top-down guidance with the bottom-up feature. Specifically, three kinds of top-down guidance factors (ie, driver intention, semantic context, and traffic rule) are integrated into our model. These factors are important for object importance estimation, but none of the existing methods simultaneously consider them. To our knowledge, this paper proposes the first on-road object importance estimation model that fuses multi-fold top-down guidance factors with bottom-up feature. Extensive experiments demonstrate that our model outperforms state-of-the-art methods by large margins, achieving 23.1% Average Precision (AP) improvement compared with the recently proposed model (ie, Goal).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。