arXiv:2607.02708cs.CV2026-07

高分辨率全局注意力能提升冻结模型的分割精度,尤其适合伪装和海洋动物识别。

When Does Resolution Help a Frozen Backbone? Global Attention at Resolution Predicts Scalable Adaptation for Camouflaged and Marine Animal Segmentation

论文配图:When Does Resolution Help a Frozen Backbone? Global Attention at Resolution Predicts Scalable Adaptation for Camouflaged and Marine Animal Segmentation
图 1 · 摘自论文原文
  • 用高分辨率全局注意力的模型更易通过低秩适配器提升性能
  • 在四个伪装动物数据集上取得最佳S-measure,海洋动物基准达mIoU 0.878
  • 适用于需要轻量微调的细粒度分割任务

将冻结的视觉基础模型适配到细粒度分割任务,很大程度上依赖骨干网络的选择。若骨干网络对高分辨率标记集应用全局注意力,则低秩适配器可将分辨率转化为准确率。各向同性ViT在全网格上全局关注,且随分辨率提升持续改进;层级化骨干网络早期注意力局限于局部窗口,并在全局阶段前下采样网格,导致性能在较低分辨率处饱和。一项包含六种骨干的受控研究验证了这一模式,修改骨干结构表明:下采样是关键因素,移除全局注意力则无此效果。该现象仅在低秩适配下显著。在固定流程SALT(Side-stem, Attention-gated U-Net, Low-rank Tuning)下,仅需一次RGB输入,强各向同性骨干在四个匹配的伪装动物数据集上获得最优S-measure,且在所有海洋与显著性数据集上领先。在两个海洋动物基准上达到新基准,MAS3K mIoU 0.878。

原文摘要 · Abstract (English)

Adapting frozen vision foundation models to fine-grained segmentation now largely depends on backbone selection. Whether the backbone applies global attention to a high-resolution token set predicts whether a low-rank adapter turns resolution into accuracy. Isotropic ViTs attend globally over the full grid and keep improving with resolution; hierarchical backbones confine early attention to local windows and pool the grid before their global stages, plateauing at lower resolutions. A controlled six-backbone study establishes the pattern, and editing the backbone points to the cause: pooling keeps the benefit, removing global attention does not. The effect is specific to low-rank adaptation. Under one fixed pipeline, SALT (Side-stem, Attention-gated U-Net, Low-rank Tuning), one RGB-only pass on a strong isotropic backbone wins the best S-measure on the four data-matched camouflaged sets, and leads every marine and salient set. It reaches a new state of the art on both marine-animal benchmarks (MAS3K mIoU 0.878).

图像分割视觉模型低秩适配海洋动物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。