arXiv:2409.20213cs.CV2024-09

通过模拟眼动聚焦提升视觉推理的泛化与样本效率

Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning

  • 基于眼动机制设计分步聚焦的视觉感知系统
  • 在多个任务上达到顶尖性能,更适应未知物体输入
  • 适合关注少样本学习与跨域泛化的研究者

人类理解视觉关系的能力远超当前AI系统,尤其在面对未见物体时。尽管AI难以判断两个新物体是否相似,人类却能轻松完成。主动视觉理论认为,视觉关系的学习依赖于眼球运动带来的注视行为,其低维空间信息有助于表征图像各部分间的关系。受此启发,我们提出一种基于瞥视(Glimpse)的主动感知系统(GAP),该系统按重要性顺序分步聚焦图像关键区域,并以高分辨率处理。系统利用瞥视动作的位置信息及其周围视觉内容,建模图像各部分间的关联。实验表明,这种机制对提取超越直接视觉内容的关系至关重要。所提方法在多个视觉推理任务中达到当前最优表现,且在样本效率和分布外泛化能力方面显著优于先前模型。

原文摘要 · Abstract (English)

Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to out-of-distribution visual inputs than prior models.

视觉推理主动感知少样本学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。