首次实现驾驶注意力的‘看哪里、看什么、为什么’三重预测,提升可解释性。
Where, What, Why: Towards Explainable Driver Attention Prediction
- 联合预测注意力位置、语义内容与认知原因,构建可解释框架。
- 在真实驾驶场景中验证,跨数据集泛化能力强,准确率显著提升。
- 适合自动驾驶、人机交互与驾驶员训练系统研发者参考。
驾驶任务中的注意力建模是自动驾驶与认知科学的核心挑战。现有方法多仅生成注视热图,难以揭示特定情境下注意力分配的认知动因,限制了对注意力机制的深层理解。为此,本文提出可解释驾驶注意力预测新范式,联合预测空间注意力区域(where)、解析关注语义(what)并提供认知推理(why)。为此构建了首个大规模可解释驾驶注意力数据集W3DA,涵盖正常、危急及事故等多样化驾驶场景,包含详尽的语义与因果标注。进一步提出基于大语言模型的LLada框架,统一像素建模、语义解析与认知推理,实现端到端建模。大量实验表明,该方法在跨数据集与驾驶条件下均具强泛化能力,为深入理解驾驶注意力机制迈出关键一步,对自动驾驶、智能驾培与人机交互具有重要价值。
原文摘要 · Abstract (English)
Modeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to capture the cognitive motivations behind attention allocation in specific contexts, which limits deeper understanding of attention mechanisms. To bridge this gap, we introduce Explainable Driver Attention Prediction, a novel task paradigm that jointly predicts spatial attention regions (where), parses attended semantics (what), and provides cognitive reasoning for attention allocation (why). To support this, we present W3DA, the first large-scale explainable driver attention dataset. It enriches existing benchmarks with detailed semantic and causal annotations across diverse driving scenarios, including normal conditions, safety-critical situations, and traffic accidents. We further propose LLada, a Large Language model-driven framework for driver attention prediction, which unifies pixel modeling, semantic parsing, and cognitive reasoning within an end-to-end architecture. Extensive experiments demonstrate the effectiveness of LLada, exhibiting robust generalization across datasets and driving conditions. This work serves as a key step toward a deeper understanding of driver attention mechanisms, with significant implications for autonomous driving, intelligent driver training, and human-computer interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。