arXiv:2609.04088cs.CV2026-09

用仿生视野聚焦关键区域,大幅降低计算量仍保持高精度语义理解。

Efficient Semantic Understanding from Digital Foveation

论文配图:Efficient Semantic Understanding from Digital Foveation
图 1 · 摘自论文原文
  • 通过视觉焦点动态选择关注区域,结合高低分辨率信息融合。
  • 单次聚焦仅需4.7%算力,达到基准95.9%的识别准确率。
  • 适合追求高效推理的部署场景,推动评估方式从像素级转向对象级。

密集语义分割对全图均匀分配计算资源,与场景复杂度和任务相关性无关。受生物视觉启发,我们探究数字福马化感知是否能更高效实现语义理解。提出一种轻量级主动视觉流程,融合显著性驱动的注视选择、高分辨率中心视区观测、低分辨率上下文信息、语义累积与自适应计算。超越传统密集预测指标,采用对象级评估衡量稀疏观测下的语义理解能力。在 ADE20K-Object 数据集上,单次福马化观测即达基准 Top-1 准确率的 95.9%,Top-3 准确率的 96.9%,计算成本仅占 4.7%。场景层面,语义累积恢复了 90.6% 的基准物体召回率,计算使用率为 58.6%。结果表明,在选择性分配计算的前提下,稀疏观测仍可产生强大语义理解,凸显主动视觉作为统一密集处理的高效替代方案,并推动评估协议向像素级之外发展。

原文摘要 · Abstract (English)

Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.

主动视觉语义分割计算效率福马化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。