受人眼视觉启发,用简单结构实现高效高精度显著物检测。
DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
- 模仿人眼双通路视觉机制,设计纯Transformer结构
- 在5个数据集上超越25种主流方法,速度提升60%
- 无需多模态数据,在伪装/水下场景也表现优异
当前显著物检测(SOD)方法聚焦于语义增强、边界优化、辅助任务监督和多模态融合四个方向,演变为复杂多阶段架构,包含专用融合模块、边缘引导学习和精细注意力机制。然而,这种复杂性导致特征冗余和跨组件干扰,抑制显著线索,陷入性能瓶颈。相比之下,人类视觉无需复杂架构即可高效识别显著目标。由此提出核心问题:能否设计一种生物启发、架构简洁的SOD框架,在保持顶尖精度的同时,兼具计算效率与可解释性?本文正面回答该问题,提出DualGazeNet——一种基于人眼视觉系统中鲁棒表征学习与镁-小细胞双通路处理机制的纯Transformer框架,引入皮层注意调制。在五个RGB-SOD基准测试中,DualGazeNet持续超越25种先进CNN与Transformer方法。平均而言,相比四种同规模变压器基线(VST++、MDSAM、Sam2unet、BiRefNet),推理速度提升约60%,浮点运算量减少53.4%。此外,DualGazeNet展现强跨域泛化能力,在无额外模态支持下,于伪装及水下显著物检测任务中取得领先或竞争力表现。
原文摘要 · Abstract (English)
Recent salient object detection (SOD) methods aim to improve performance in four key directions: semantic enhancement, boundary refinement, auxiliary task supervision, and multi-modal fusion. In pursuit of continuous gains, these approaches have evolved toward increasingly sophisticated architectures with multi-stage pipelines, specialized fusion modules, edge-guided learning, and elaborate attention mechanisms. However, this complexity paradoxically introduces feature redundancy and cross-component interference that obscure salient cues, ultimately reaching performance bottlenecks. In contrast, human vision achieves efficient salient object identification without such architectural complexity. This contrast raises a fundamental question: can we design a biologically grounded yet architecturally simple SOD framework that dispenses with most of this engineering complexity, while achieving state-of-the-art accuracy, computational efficiency, and interpretability? In this work, we answer this question affirmatively by introducing DualGazeNet, a biologically inspired pure Transformer framework that models the dual biological principles of robust representation learning and magnocellular-parvocellular dual-pathway processing with cortical attention modulation in the human visual system. Extensive experiments on five RGB SOD benchmarks show that DualGazeNet consistently surpasses 25 state-of-the-art CNN- and Transformer-based methods. On average, DualGazeNet achieves about 60\% higher inference speed and 53.4\% fewer FLOPs than four Transformer-based baselines of similar capacity (VST++, MDSAM, Sam2unet, and BiRefNet). Moreover, DualGazeNet exhibits strong cross-domain generalization, achieving leading or highly competitive performance on camouflaged and underwater SOD benchmarks without relying on additional modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。