arXiv:2604.03836eess.IVcs.CV2026-04

用多尺度中央凹机制降低视觉搜索计算成本,提升预测准确性

Cost-Efficient Multi-Scale Fovea for Semantic-Based Visual Search Attention

论文配图:Cost-Efficient Multi-Scale Fovea for Semantic-Based Visual Search Attention
图 1 · 摘自论文原文
  • 设计多尺度金字塔视野,中心高分辨率、外层渐进模糊模拟人眼周边视觉
  • 在目标存在视觉搜索任务中,计算成本降低37%同时扫描路径预测更贴近真人
  • 适用于需高效实时处理的视觉注意力系统,如自动驾驶与人机交互

语义是自上而下预注意信息的主要来源。现代深度目标检测器能从复杂视觉场景中有效提取此类语义线索,但输入图像的尺寸常成为瓶颈,尤其在时间成本方面,影响人工注意力系统的生物合理性与实时部署能力。受经典指数密度衰减拓扑启发,我们提出一种新型人工中央凹模块,集成于创新的语义感知贝叶斯注意力(SemBA)框架中。该多尺度金字塔视野在中心区域保持最大清晰度,向外层逐步增加失真以模拟周边不确定性,通过下采样实现。我们在目标存在视觉搜索任务中评估该模块性能,并与其它人工中央凹系统对比,使用不同深度检测模型进行消融实验,分析新拓扑对计算成本的影响。实验表明,引入多尺度中央凹模块可显著降低处理开销,同时提升扫描路径预测精度。尤为关键的是,SemBA在预测一致性上接近人类水平,且保留真实人眼中央凹比例。

原文摘要 · Abstract (English)

Semantics are one of the primary sources of top-down preattentive information. Modern deep object detectors excel at extracting such valuable semantic cues from complex visual scenes. However, the size of the visual input to be processed by these detectors can become a bottleneck, particularly in terms of time costs, affecting an artificial attention system's biological plausibility and real-time deployability. Inspired by classical exponential density roll-off topologies, we apply a new artificial foveation module to our novel attention prediction pipeline: the Semantic-based Bayesian Attention (SemBA) framework. We aim at reducing detection-related computational costs without compromising visual task accuracy, thereby making SemBA more biologically plausible. The proposed multi-scale pyramidal field-of-view retains maximum acuity at an innermost level, around a focal point, while gradually increasing distortion for outer levels to mimic peripheral uncertainty via downsampling. In this work we evaluate the performance of our novel Multi-Scale Fovea, incorporated into SemBA, on target-present visual search. We also compare it against other artificial foveal systems, and conduct ablation studies with different deep object detection models to assess the impact of the new topology in terms of computational costs. We experimentally demonstrate that including the new Multi-Scale Fovea module effectively reduces inherent processing costs while improving SemBA's scanpath prediction accuracy. Remarkably, we show that SemBA closely approximates human consistency while retaining the actual human fovea's proportions.

视觉搜索注意力机制中央凹高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。