arXiv:2411.13302cs.CV2024-11被引 2

给行人意图预测加解释,让AI决策更可懂。

Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

  • 用跨模态学习同时预测行人意图和背后原因。
  • 在PIE++数据集上准确率提升5.6%,F1提升7%。
  • 解释内容可帮助理解行为,适合自动驾驶安全研究。

随着自动驾驶系统对道路安全的重视,保护弱势道路使用者(如行人)的需求日益迫切。行人意图预测是其中一项挑战性任务,现有方法通常融合视觉与运动特征进行二元分类(过/不过)。然而,目前尚无工作将预测结果与人类可理解的原因结合。本文提出一种新问题设定:探索行人行为背后的直观理由。为此,构建了多标签文本解释的PIE++数据集,并提出名为MINDREAD的多任务学习框架,通过跨模态表示学习同时预测意图及其原因。实验表明,在PIE++数据集上,意图预测的准确率提升5.6%,F1-score提升7%;在常用JAAD数据集上准确率也提升4.4%。定量与定性评估及用户研究均验证了该方法的有效性。

原文摘要 · Abstract (English)

With the increased importance of autonomous navigation systems has come an increasing need to protect the safety of Vulnerable Road Users (VRUs) such as pedestrians. Predicting pedestrian intent is one such challenging task, where prior work predicts the binary cross/no-cross intention with a fusion of visual and motion features. However, there has been no effort so far to hedge such predictions with human-understandable reasons. We address this issue by introducing a novel problem setting of exploring the intuitive reasoning behind a pedestrian's intent. In particular, we show that predicting the 'WHY' can be very useful in understanding the 'WHAT'. To this end, we propose a novel, reason-enriched PIE++ dataset consisting of multi-label textual explanations/reasons for pedestrian intent. We also introduce a novel multi-task learning framework called MINDREAD, which leverages a cross-modal representation learning framework for predicting pedestrian intent as well as the reason behind the intent. Our comprehensive experiments show significant improvement of 5.6% and 7% in accuracy and F1-score for the task of intent prediction on the PIE++ dataset using MINDREAD. We also achieved a 4.4% improvement in accuracy on a commonly used JAAD dataset. Extensive evaluation using quantitative/qualitative metrics and user studies shows the effectiveness of our approach.

意图预测跨模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。