通过动态聚焦关键区域,实现高效任务处理的神经网络架构
Task-Driven Fixation Network: An Efficient Architecture with Fixation Selection
- 双通道结构:低分辨率全局特征 + 高分辨率局部特征
- 任务驱动生成焦点点,减少计算量并保持性能
- 适合需要实时处理的视觉任务,如目标检测与分割
本文提出一种新型神经网络架构,具备自动聚焦点选择能力,可高效完成复杂任务,同时减少模型规模和计算开销。该模型包含三个部分:低分辨率通道,用于提取输入图像的全局特征;高分辨率通道,依次提取局部细节特征;以及融合两通道特征的混合编码模块。该模块的核心是焦点点生成器,能以任务为导向动态生成关注区域,使高分辨率通道集中于感兴趣区域。此方法避免对整图进行高分辨率遍历分析,在保持任务性能的同时显著提升效率。
原文摘要 · Abstract (English)
This paper presents a novel neural network architecture featuring automatic fixation point selection, designed to efficiently address complex tasks with reduced network size and computational overhead. The proposed model consists of: a low-resolution channel that captures low-resolution global features from input images; a high-resolution channel that sequentially extracts localized high-resolution features; and a hybrid encoding module that integrates the features from both channels. A defining characteristic of the hybrid encoding module is the inclusion of a fixation point generator, which dynamically produces fixation points, enabling the high-resolution channel to focus on regions of interest. The fixation points are generated in a task-driven manner, enabling the automatic selection of regions of interest. This approach avoids exhaustive high-resolution analysis of the entire image, maintaining task performance and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。