arXiv:2506.11136cs.CVeess.IV2025-06NeurIPS被引 26

JAFAR可任意提升视觉模型特征分辨率,无需高分辨率监督

JAFAR: Jack up Any Feature at Any Resolution

  • 用注意力机制融合低层图像特征与高层语义特征
  • 在低倍率训练下仍能良好泛化至高分辨率输出
  • 适合需要精细空间细节的下游视觉任务

基础视觉编码器在密集视觉任务中至关重要,但其输出的空间特征分辨率较低,需通过特征上采样生成下游任务所需的高分辨率表示。本文提出JAFAR,一种轻量且灵活的特征上采样器,可将任意基础视觉编码器的特征提升至任意目标分辨率。JAFAR采用基于注意力的模块,利用空间特征变换(SFT)调制,促进由低层图像特征生成的高分辨率查询与语义丰富的低分辨率键之间的语义对齐。值得注意的是,尽管缺乏高分辨率监督,我们证明在低上采样比率和分辨率下学习,能显著泛化至更高输出尺度。大量实验表明,JAFAR能有效恢复细粒度空间细节,并在多种下游任务中持续优于现有特征上采样方法。

原文摘要 · Abstract (English)

Foundation Vision Encoders have become essential for a wide range of dense vision tasks. However, their low-resolution spatial feature outputs necessitate feature upsampling to produce the high-resolution modalities required for downstream tasks. In this work, we introduce JAFAR, a lightweight and flexible feature upsampler that enhances the spatial resolution of visual features from any Foundation Vision Encoder to an arbitrary target resolution. JAFAR employs an attention-based module designed to promote semantic alignment between high-resolution queries, derived from low-level image features, and semantically enriched low-resolution keys, using Spatial Feature Transform (SFT) modulation. Notably, despite the absence of high-resolution supervision, we demonstrate that learning at low upsampling ratios and resolutions generalizes remarkably well to significantly higher output scales. Extensive experiments show that JAFAR effectively recovers fine-grained spatial details and consistently outperforms existing feature upsampling methods across a diverse set of downstream tasks. Project page at https://jafar-upsampler.github.io

特征上采样视觉编码器空间对齐轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。