arXiv:2503.16910cs.CV2025-03被引 5

针对交通场景设计首个大规模显著性检测数据集,提升自动驾驶安全感知能力。

Salient Object Detection in Traffic Scene through the TSOD10K Dataset

  • 基于Mamba架构设计新型Tramba模型,融合双频视觉状态空间与驾驶注意力机制。
  • 在TSOD10K数据集上实现mAP 78.3,显著优于22个自然场景模型。
  • 适用于智能网联汽车、驾驶辅助系统等交通安全关键场景。

交通显著性检测(TSOD)旨在通过结合语义重要性(如碰撞风险)与视觉显著性,分割对行车安全至关重要的目标。与自然场景显著性检测(NSI-SOD)侧重视觉突出区域不同,TSOD更关注因语义风险而需立即注意的目标,即使其视觉对比度较低。为应对缺乏任务专用基准的问题,本文构建首个大规模像素级标注的TSOD数据集——TSOD10K,涵盖多种真实交通场景下的多样化目标类别及复杂天气/光照条件(如雾天、雪灾、低对比度、弱光)。方法上,提出基于Mamba的Tramba模型:引入双频视觉状态空间模块,通过移位窗口划分与扩张扫描,分层解耦高低频成分以增强细节与全局结构感知;设计面向交通的螺旋形2D选择性扫描(Helix-SS2D)机制,在捕捉多方向全局空间依赖的同时注入驾驶注意力先验。在TSOD10K上评估Tramba及22个现有NSI-SOD模型,验证其优越性能(mAP 78.3),为智能交通系统的安全感知分析奠定首个基础。

原文摘要 · Abstract (English)

Traffic Salient Object Detection (TSOD) aims to segment the objects critical to driving safety by combining semantic (e.g., collision risks) and visual saliency. Unlike SOD in natural scene images (NSI-SOD), which prioritizes visually distinctive regions, TSOD emphasizes the objects that demand immediate driver attention due to their semantic impact, even with low visual contrast. This dual criterion, i.e., bridging perception and contextual risk, re-defines saliency for autonomous and assisted driving systems. To address the lack of task-specific benchmarks, we collect the first large-scale TSOD dataset with pixel-wise saliency annotations, named TSOD10K. TSOD10K covers the diverse object categories in various real-world traffic scenes under various challenging weather/illumination variations (e.g., fog, snowstorms, low-contrast, and low-light). Methodologically, we propose a Mamba-based TSOD model, termed Tramba. Considering the challenge of distinguishing inconspicuous visual information from complex traffic backgrounds, Tramba introduces a novel Dual-Frequency Visual State Space module equipped with shifted window partitioning and dilated scanning to enhance the perception of fine details and global structure by hierarchically decomposing high/low-frequency components. To emphasize critical regions in traffic scenes, we propose a traffic-oriented Helix 2D-Selective-Scan (Helix-SS2D) mechanism that injects driving attention priors while effectively capturing global multi-direction spatial dependencies. We establish a comprehensive benchmark by evaluating Tramba and 22 existing NSI-SOD models on TSOD10K, demonstrating Tramba's superiority. Our research establishes the first foundation for safety-aware saliency analysis in intelligent transportation systems.

交通场景显著性检测Mamba自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。