arXiv:2505.01249cs.CVcs.LG2025-05

用线性视网膜变换融合多注视点,提升视觉信息整合效率

Fusing Foveal Fixations Using Linear Retinal Transformations and Bayesian Experimental Design

  • 将注视点的视网膜变换建模为线性下采样,实现精确推断
  • 在Frey人脸和MNIST数据集上验证了模型的有效性
  • 结合贝叶斯实验设计,优化下一步注视位置选择

人类(及许多脊椎动物)需将多个场景注视点的信息融合,形成完整表征,每个注视点使用高分辨率中央凹和逐渐降低的周边分辨率。本文将注视点的视网膜变换显式建模为对场景高分辨率潜在图像的线性下采样,利用已知几何结构。该线性变换使我们能够在因子分析(FA)及其混合模型中进行精确推断。此外,该方法可将“下一步看哪里”的决策问题建模为基于期望信息增益准则的贝叶斯实验设计问题。在Frey faces和MNIST数据集上的实验验证了模型的有效性。

原文摘要 · Abstract (English)

Humans (and many vertebrates) face the problem of fusing together multiple fixations of a scene in order to obtain a representation of the whole, where each fixation uses a high-resolution fovea and decreasing resolution in the periphery. In this paper we explicitly represent the retinal transformation of a fixation as a linear downsampling of a high-resolution latent image of the scene, exploiting the known geometry. This linear transformation allows us to carry out exact inference for the latent variables in factor analysis (FA) and mixtures of FA models of the scene. Further, this allows us to formulate and solve the choice of "where to look next" as a Bayesian experimental design problem using the Expected Information Gain criterion. Experiments on the Frey faces and MNIST datasets demonstrate the effectiveness of our models.

视觉融合贝叶斯设计因子分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。