arXiv:2409.09446cs.CVcs.AI2024-09

通过多模态概念提升行人动作预测可解释性,兼顾准确与透明。

MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction

  • 用多模态概念整合预测,实现跨模态关联解释
  • 引入局部注意力机制,捕捉输入细节特征
  • 通过正则化避免模式坍缩,提升泛化能力

行人动作预测在自动驾驶等场景中具有重要意义,但现有方法缺乏可解释性。本文提出新框架MulCPred,基于训练样本中的多模态概念进行预测解释。针对已有概念方法的三大局限——无法直接处理多模态、缺乏局部关注能力、易出现模式坍缩——提出三种改进:1)线性聚合器融合多模态概念激活结果,提供预测前的可解释性;2)通道重校准模块聚焦局部时空区域,增强概念局部性;3)特征正则化损失促使概念学习多样化模式。在多个数据集和任务上评估显示,MulCPred在不明显降低性能的前提下显著提升可解释性。进一步移除不可识别概念后,跨数据集预测性能提升,表明其具备良好的泛化潜力。

原文摘要 · Abstract (English)

Pedestrian action prediction is of great significance for many applications such as autonomous driving. However, state-of-the-art methods lack explainability to make trustworthy predictions. In this paper, a novel framework called MulCPred is proposed that explains its predictions based on multi-modal concepts represented by training samples. Previous concept-based methods have limitations including: 1) they cannot directly apply to multi-modal cases; 2) they lack locality to attend to details in the inputs; 3) they suffer from mode collapse. These limitations are tackled accordingly through the following approaches: 1) a linear aggregator to integrate the activation results of the concepts into predictions, which associates concepts of different modalities and provides ante-hoc explanations of the relevance between the concepts and the predictions; 2) a channel-wise recalibration module that attends to local spatiotemporal regions, which enables the concepts with locality; 3) a feature regularization loss that encourages the concepts to learn diverse patterns. MulCPred is evaluated on multiple datasets and tasks. Both qualitative and quantitative results demonstrate that MulCPred is promising in improving the explainability of pedestrian action prediction without obvious performance degradation. Furthermore, by removing unrecognizable concepts from MulCPred, the cross-dataset prediction performance is improved, indicating the feasibility of further generalizability of MulCPred.

动作预测可解释性多模态行人分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。