用扩散模型预测驾驶员视觉注意力,提升智能驾驶安全感知
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
- 将注意力预测建模为条件扩散去噪过程,融合多尺度特征
- 在四个公开数据集上达到当前最优性能,显著超越基线方法
- 结合大模型增强语义推理,适合智能座舱与风险感知场景
驾驶员的视觉注意力为预判潜在危险提供关键线索,直接影响决策与操控行为,其缺失会危及交通安全。为模拟驾驶员感知模式并推动智能车辆中的视觉注意力预测,我们提出DiffAttn,一种基于扩散模型的框架,将该任务建模为条件扩散-去噪过程,从而更准确地捕捉驾驶员注意力分布。为同时捕获局部与全局场景特征,采用Swin Transformer作为编码器,并设计解码器,结合特征融合金字塔实现跨层交互,辅以密集、多尺度的条件扩散机制,共同增强去噪学习并建模精细的局部与全局场景上下文。此外,引入大语言模型(LLM)层以增强自上而下的语义推理能力,提升对安全关键线索的敏感性。在四个公开数据集上的大量实验表明,DiffAttn达到当前最优(SoTA)性能,超越多数基于视频、基于自上而下特征及基于LLM增强的基线方法。该框架还支持可解释的以驾驶员为中心的场景理解,具备提升智能车辆中人机交互、风险感知与驾驶员状态监测的潜力。
原文摘要 · Abstract (English)
Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety. To emulate drivers' perception patterns and advance visual attention prediction for intelligent vehicles, we propose DiffAttn, a diffusion-based framework that formulates this task as a conditional diffusion-denoising process, enabling more accurate modeling of drivers' attention. To capture both local and global scene features, we adopt Swin Transformer as encoder and design a decoder that combines a Feature Fusion Pyramid for cross-layer interaction with dense, multi-scale conditional diffusion to jointly enhance denoising learning and model fine-grained local and global scene contexts. Additionally, a large language model (LLM) layer is incorporated to enhance top-down semantic reasoning and improve sensitivity to safety-critical cues. Extensive experiments on four public datasets demonstrate that DiffAttn achieves state-of-the-art (SoTA) performance, surpassing most video-based, top-down-feature-driven, and LLM-enhanced baselines. Our framework further supports interpretable driver-centric scene understanding and has the potential to improve in-cabin human-machine interaction, risk perception, and drivers' state measurement in intelligent vehicles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。