arXiv:2510.12764cs.CVcs.LG2025-10中稿 · ICLR被引 27

AnyUp无需训练即可通用提升任意视觉特征分辨率

AnyUp: Universal Feature Upsampling

  • 设计可泛化至不同特征提取器的推理时通用上采样架构
  • 在多种特征类型上实现当前最优上采样效果
  • 适合需要快速部署上采样模块的研究者和工程师

我们提出AnyUp,一种可在任意视觉特征和任意分辨率下应用的特征上采样方法,无需针对特定编码器进行训练。现有基于学习的上采样器(如DINO或CLIP特征)需为每个特征提取器重新训练,无法在推理时泛化到不同特征类型。本文提出一种推理时特征无关的上采样架构,缓解该限制并提升上采样质量。实验表明,AnyUp在上采样特征上达到新最优性能,可泛化至多种特征类型,并在保持特征语义的同时具备高效性与广泛适用性。

原文摘要 · Abstract (English)

We introduce AnyUp, a method for feature upsampling that can be applied to any vision feature at any resolution, without encoder-specific training. Existing learning-based upsamplers for features like DINO or CLIP need to be re-trained for every feature extractor and thus do not generalize to different feature types at inference time. In this work, we propose an inference-time feature-agnostic upsampling architecture to alleviate this limitation and improve upsampling quality. In our experiments, AnyUp sets a new state of the art for upsampled features, generalizes to different feature types, and preserves feature semantics while being efficient and easy to apply to a wide range of downstream tasks.

特征上采样视觉模型通用架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。