提出分层注意力模型,让机器视觉更像人眼
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
- 分两层递归结构,分别负责定位和任务执行
- 在图像分类任务中准确率优于CNN、RAM和DRAM
- 自动生成类人眼的注视与快速转移行为
受中央视觉启发,硬注意力模型具有可解释性和参数高效性。然而,现有模型如递归视觉注意力模型(RAM)和深度递归注意力模型(DRAM)未能建模人类视觉系统的层次结构,导致视觉探索行为失衡,注意力要么过于停留,要么频繁跳跃,偏离真实眼动模式。本文提出多层级递归注意力模型(MRAM),显式模拟人类视觉处理的神经层次结构。通过将瞥视位置生成与任务执行分离到两个递归层,MRAM自然涌现出注视与扫视之间的平衡行为。实验表明,MRAM不仅实现了更类人的注意力动态,还在标准图像分类基准上持续优于CNN、RAM和DRAM基线模型。
原文摘要 · Abstract (English)
Inspired by foveal vision, hard attention models promise interpretability and parameter economy. However, existing models like the Recurrent Model of Visual Attention (RAM) and Deep Recurrent Attention Model (DRAM) failed to model the hierarchy of human vision system, that compromise on the visual exploration dynamics. As a result, they tend to produce attention that are either overly fixational or excessively saccadic, diverging from human eye movement behavior. In this paper, we propose a Multi-Level Recurrent Attention Model (MRAM), a novel hard attention framework that explicitly models the neural hierarchy of human visual processing. By decoupling the function of glimpse location generation and task execution in two recurrent layers, MRAM emergent a balanced behavior between fixation and saccadic movement. Our results show that MRAM not only achieves more human-like attention dynamics, but also consistently outperforms CNN, RAM and DRAM baselines on standard image classification benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。