arXiv:2410.11688cs.CVcs.NE2024-10被引 1

用眼球聚焦模拟提升视网膜假体视觉感知,准确率达87.72%。

Visual Fixation-Based Retinal Prosthetic Simulation

  • 基于视觉聚焦机制,用ViT注意力图定位关键图像区域
  • 在ImageNet子集上实现87.72%分类准确率,远超下采样法的40.59%
  • 适合视网膜假体研发与低分辨率视觉重建研究者

本研究提出一种基于视觉聚焦的视网膜假体仿真框架,借鉴眼跳机制,通过视觉Transformer的自注意力图预测显著图像块,模拟视觉聚焦。这些图像块经可训练U-Net编码后,利用pulse2percept框架生成视觉幻象。通过引入可学习编码器,优化传递给视网膜植入物的视觉信息,缓解电极阵列分辨率有限及输入刺激与光点感知间的畸变问题。预测的视觉感知使用自监督DINOv2基础模型评估,可选可学习线性层以提升分类精度。在基于真实受试者生理数据设定计算参数的ImageNet验证集子集上,该方法达到87.72%分类准确率,显著优于下采样法的40.59%,接近健康上限92.76%。结果表明该方法有望在有限分辨率下生成更具语义理解性的感知。

原文摘要 · Abstract (English)

This study proposes a retinal prosthetic simulation framework driven by visual fixations, inspired by the saccade mechanism, and assesses performance improvements through end-to-end optimization in a classification task. Salient patches are predicted from input images using the self-attention map of a vision transformer to mimic visual fixations. These patches are then encoded by a trainable U-Net and simulated using the pulse2percept framework to predict visual percepts. By incorporating a learnable encoder, we aim to optimize the visual information transmitted to the retinal implant, addressing both the limited resolution of the electrode array and the distortion between the input stimuli and resulting phosphenes. The predicted percepts are evaluated using the self-supervised DINOv2 foundation model, with an optional learnable linear layer for classification accuracy. On a subset of the ImageNet validation set, the fixation-based framework achieves a classification accuracy of 87.72%, using computational parameters based on a real subject's physiological data, significantly outperforming the downsampling-based accuracy of 40.59% and approaching the healthy upper bound of 92.76%. Our approach shows promising potential for producing more semantically understandable percepts with the limited resolution available in retinal prosthetics.

视网膜假体视觉聚焦图像重建Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。