arXiv:2511.02652cs.CV2025-11NeurIPS被引 3

提出可微分的分层视觉分词方法,让模型自适应图像内容。

Differentiable Hierarchical Visual Tokenization

  • 基于信息准则的分层模型选择,实现像素级自适应分词。
  • 在图像分类与密集预测任务中表现媲美固定分块方法。
  • 支持直接将位图转为矢量图,适合模型改造与图形生成。

视觉变换器依赖固定的图像块令牌,忽略了图像的空间与语义结构。本文提出一种端到端可微分的分词器,能够以像素级粒度自适应图像内容,同时保持与现有架构的向后兼容性,便于对预训练模型进行改造。该方法采用分层模型选择结合信息准则,在图像级分类与密集预测任务中均取得具有竞争力的性能,甚至支持开箱即用的位图到矢量图转换。

原文摘要 · Abstract (English)

Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adapts to image content with pixel-level granularity while remaining backward-compatible with existing architectures for retrofitting pretrained models. Our method uses hierarchical model selection with information criteria to provide competitive performance in both image-level classification and dense-prediction tasks, and even supports out-of-the-box raster-to-vector conversion.

视觉分词可微分视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。