arXiv:2607.10066eess.IV2026-07

基于生物视觉的类脑系统,193微秒完成视觉任务,精度提升超三成。

A neuromorphic vision system for open-world visual intelligence

  • 通过光场选择与目标预判逐步提炼关键信息
  • 在8个开放场景中追踪、分割、预测精度分别提升25.54%~36.10%
  • 延迟降低30.6倍,适合实时嵌入式视觉应用

在非结构化开放世界环境中,高效且鲁棒的视觉智能仍是重大挑战,现有方法常依赖计算密集型神经架构或专用传感器,通用性受限。受生物视觉与信息瓶颈理论启发,我们提出一种类脑视觉系统,通过硬件实现的任务牵引机制(task traction mechanism)进行面向任务的信息蒸馏。该系统结合偏振敏感成像仪与阻变存储器(RRAM)阵列,通过光场选择、兴趣区域提取和目标预判,逐级过滤冗余信息。系统执行时间仅193 μs。在八个复杂开放世界场景中的评估显示,目标追踪、分割与轨迹预测准确率分别提升25.54%、37.73%和36.10%,平均延迟较最先进方案降低30.6倍。

原文摘要 · Abstract (English)

Time-efficient and robust visual intelligence remains a critical challenge in unstructured open-world environments, yet current approaches often rely on computationally intensive neural architectures or task-specific sensors with limited versatility. Inspired by biological vision and information bottleneck theory, we report a neuromorphic vision system that performs task-oriented visual intelligence through an information distillation strategy (named as task traction mechanism) implemented on hardware. The system integrates a polarization-sensitive imager with a resistive random-access memory (RRAM) array to progressively distill task-relevant information via light field selection, region of interest extraction, and target anticipation. The neuromorphic vision system conducts visual tasks within an execution time of 193 μs. Evaluation across eight challenging open-world scenarios shows accuracy improvements of 25.54%, 37.73%, and 36.10% for object tracking, object segmentation, and trajectory prediction, respectively, together with an average 30.6-fold reduction in latency relative to state-of-the-art solutions.

类脑计算视觉感知低延迟神经形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。