用自适应超像素编码提升视觉模型对图像内容的感知能力
Representation Learning with Adaptive Superpixel Coding
- 基于Transformer设计自适应超像素分块,动态匹配图像内容
- 在标准图像下游任务中表现优于主流方法
- 适合需要灵活图像结构建模的研究与应用
深度学习视觉模型通常针对特定模态设计,且依赖领域特定假设,如几乎所有现有视觉模型采用网格结构。本文提出一种基于Transformer的自监督模型——自适应超像素编码(Adaptive Superpixel Coding, ASC)。其核心思想是克服传统Vision Transformers对固定尺寸、非自适应图像块划分的依赖。ASC采用能动态适应图像内容的自适应超像素层。我们分析了该方法的关键特性,发现其在标准图像下游任务基准测试中优于广泛使用的替代方案。
原文摘要 · Abstract (English)
Deep learning vision models are typically tailored for specific modalities and often rely on domain-specific assumptions, such as the grid structures used by nearly all existing vision models. In this work, we propose a self-supervised model based on Transformers, which we call Adaptive Superpixel Coding (ASC). The key insight of our model is to overcome the limitations of traditional Vision Transformers, which depend on fixed-size and non-adaptive patch partitioning. Instead, ASC employs adaptive superpixel layers that dynamically adjust to the underlying image content. We analyze key properties of the approach that make it effective, and find that our method outperforms widely-used alternatives on standard image downstream task benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。