用重建差异引导注意力,提升高分辨率图像的注视点精度。
Advancing TDFN: Precise Fixation Point Generation Using Reconstruction Differences
- 基于输入与重建图像的差异生成注视点,替代强化学习训练。
- 在高分辨率图像上实现像素级精确注视点,分类准确率显著提升。
- 适合需要精准视觉注意力的模型优化,如图像识别任务。
Wang 和 Wang (2025) 提出基于注视机制的任务驱动注视网络(TDFN),利用低分辨率信息与注视点附近的高分辨率细节完成特定视觉任务,并采用强化学习生成注视点。然而,强化学习在高分辨率图像上生成像素级精确注视点时训练困难。本文提出一种改进方法,通过输入图像与重建图像之间的差异来训练注视点生成器,使注视点聚焦于差异显著区域。实验表明,该方法能生成高度精确的注视点,显著提升网络分类准确率,并将达到预设准确率所需的平均注视次数减少。
原文摘要 · Abstract (English)
Wang and Wang (2025) proposed the Task-Driven Fixation Network (TDFN) based on the fixation mechanism, which leverages low-resolution information along with high-resolution details near fixation points to accomplish specific visual tasks. The model employs reinforcement learning to generate fixation points. However, training reinforcement learning models is challenging, particularly when aiming to generate pixel-level accurate fixation points on high-resolution images. This paper introduces an improved fixation point generation method by leveraging the difference between the reconstructed image and the input image to train the fixation point generator. This approach directs fixation points to areas with significant differences between the reconstructed and input images. Experimental results demonstrate that this method achieves highly accurate fixation points, significantly enhances the network's classification accuracy, and reduces the average number of required fixations to achieve a predefined accuracy level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。