arXiv:2409.11923cs.CV2024-09ECCV被引 18

一种无需训练的分组聚类方法,显著提升视觉任务效率

Agglomerative Token Clustering

  • 自底向上聚类合并令牌,无额外可学习参数
  • 低保留率下仍保持顶尖性能,无需微调即达最优
  • 适合追求高效推理的视觉模型部署场景

我们提出一种新型令牌合并方法——凝聚式令牌聚类(ATC),在图像分类、图像生成及目标检测与分割任务中均持续优于以往的令牌合并与剪枝方法。ATC通过自底向上的层次聚类方式合并令牌,不引入额外可学习参数。实验发现,ATC在所有任务上均达到当前最优性能,且在未微调的情况下即可与先前最优方法相媲美。该方法在低保留率条件下尤为有效,此时仅保留少量令牌,仍能保持任务性能,极具挑战性。

原文摘要 · Abstract (English)

We present Agglomerative Token Clustering (ATC), a novel token merging method that consistently outperforms previous token merging and pruning methods across image classification, image synthesis, and object detection & segmentation tasks. ATC merges clusters through bottom-up hierarchical clustering, without the introduction of extra learnable parameters. We find that ATC achieves state-of-the-art performance across all tasks, and can even perform on par with prior state-of-the-art when applied off-the-shelf, i.e. without fine-tuning. ATC is particularly effective when applied with low keep rates, where only a small fraction of tokens are kept and retaining task performance is especially difficult.

令牌合并聚类视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。