基于生物视觉机制,用脉冲网络实现动态物体感知与注意力聚焦。
Wandering around: A bioinspired approach to visual attention through object motion sensitivity
- 通过脉冲卷积网络模拟眼球微动,利用事件相机捕捉运动物体。
- 多物体运动分割平均交并比达82.2%,低光场景目标检测准确率超89%。
- 无需训练、实时响应仅0.12秒,适合机器人等实时系统部署。
主动视觉可实现动态感知,替代依赖大规模数据和高算力的静态前馈架构。生物选择性注意机制使智能体聚焦显著区域(ROIs),在保持实时响应的同时降低计算负担。事件相机模仿哺乳动物视网膜,通过异步记录场景变化,实现高效低延迟处理。为在相机运动时区分移动物体,需具备物体运动分割能力以精准定位目标并使其居中于视野(中央凹)。将事件传感器与类脑算法结合,采用脉冲神经网络并行计算,适应动态环境。本文提出一种受生物启发的脉冲卷积神经网络注意力系统,通过动态视觉传感器(DVS)在Speck类脑硬件上模拟眼动生成事件,实现目标识别与眼跳转向。系统在理想条纹测试中表现良好,并在事件相机运动分割数据集上达到82.2%的平均交并比(mIoU)和96%的平均结构相似性(SSIM)。在办公室场景中目标检测准确率达88.8%,在低光条件下对事件辅助低光视频目标分割数据集的准确率为89.8%。实时演示显示其对动态场景响应仅需0.12秒。该系统无学习过程,鲁棒性强,适用于各类感知场景,是复杂类脑架构的可靠基础。
原文摘要 · Abstract (English)
Active vision enables dynamic visual perception, offering an alternative to static feedforward architectures in computer vision, which rely on large datasets and high computational resources. Biological selective attention mechanisms allow agents to focus on salient Regions of Interest (ROIs), reducing computational demand while maintaining real-time responsiveness. Event-based cameras, inspired by the mammalian retina, enhance this capability by capturing asynchronous scene changes enabling efficient low-latency processing. To distinguish moving objects while the event-based camera is in motion the agent requires an object motion segmentation mechanism to accurately detect targets and center them in the visual field (fovea). Integrating event-based sensors with neuromorphic algorithms represents a paradigm shift, using Spiking Neural Networks to parallelize computation and adapt to dynamic environments. This work presents a Spiking Convolutional Neural Network bioinspired attention system for selective attention through object motion sensitivity. The system generates events via fixational eye movements using a Dynamic Vision Sensor integrated into the Speck neuromorphic hardware, mounted on a Pan-Tilt unit, to identify the ROI and saccade toward it. The system, characterized using ideal gratings and benchmarked against the Event Camera Motion Segmentation Dataset, reaches a mean IoU of 82.2% and a mean SSIM of 96% in multi-object motion segmentation. The detection of salient objects reaches 88.8% accuracy in office scenarios and 89.8% in low-light conditions on the Event-Assisted Low-Light Video Object Segmentation Dataset. A real-time demonstrator shows the system's 0.12 s response to dynamic scenes. Its learning-free design ensures robustness across perceptual scenes, making it a reliable foundation for real-time robotic applications serving as a basis for more complex architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。