arXiv:2509.05606cs.CV2025-09

提出无参数的局部核对齐方法,提升视觉模型细粒度表征能力。

Patch-Level Kernel Alignment for Dense Self-Supervised Learning

  • 基于核函数的非参数对齐,捕捉高维特征分布的统计依赖关系
  • 仅用14小时单卡微调,在多个密集视觉任务上达到顶尖性能
  • 适用于预训练模型的轻量级后训练,适合快速部署和增强

密集自监督学习(Dense SSL)在提升视觉模型细粒度语义理解方面表现出色。然而,现有方法常依赖参数假设或复杂后处理(如聚类、排序),限制了灵活性与稳定性。为此,我们提出一种非参数、基于核函数的局部核对齐(PaKA)方法,通过后(前)训练优化预训练视觉编码器的密集表示。该方法设计了一个稳健有效的对齐目标,能够捕捉高维密集特征分布的内在结构。此外,我们重新审视图像级SSL继承的增强策略,提出一种适配密集自监督学习的精细化增强方案。我们的框架在预训练模型基础上进行轻量级后训练,仅需单张GPU运行14小时,即可在多个密集视觉基准上实现领先性能,充分验证其高效性与有效性。

原文摘要 · Abstract (English)

Dense self-supervised learning (SSL) methods showed its effectiveness in enhancing the fine-grained semantic understandings of vision models. However, existing approaches often rely on parametric assumptions or complex post-processing (e.g., clustering, sorting), limiting their flexibility and stability. To overcome these limitations, we introduce Patch-level Kernel Alignment (PaKA), a non-parametric, kernel-based approach that improves the dense representations of pretrained vision encoders with a post-(pre)training. Our method propose a robust and effective alignment objective that captures statistical dependencies which matches the intrinsic structure of high-dimensional dense feature distributions. In addition, we revisit the augmentation strategies inherited from image-level SSL and propose a refined augmentation strategy for dense SSL. Our framework improves dense representations by conducting a lightweight post-training stage on top of a pretrained model. With only 14 hours of additional training on a single GPU, our method achieves state-of-the-art performance across a range of dense vision benchmarks, demonstrating both efficiency and effectiveness.

自监督学习密集表征核方法后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。