arXiv:2508.16317cs.CV2025-08被引 3

视觉编码器应随任务动态调整计算量,而非依赖图像尺寸。

Vision encoders should be image size agnostic and task driven

  • 根据任务需求动态调节计算复杂度,突破图像尺寸限制。
  • 在图像分类任务中验证了方法可行性,展现潜力。
  • 受生物视觉效率启发,适合需要节能的实时视觉系统。

本文主张下一代视觉编码器应具备图像尺寸无关性与任务驱动性。灵感源自生物视觉的行为效率:人类与动物面对海量视觉数据时,会根据任务智能分配有限能量。然而现代视觉编码器并未体现这种灵活性。我们提出,视觉编码器应是动态的,其计算复杂度应由任务决定,而非图像大小。为此,我们提供了一个图像分类任务的可行性验证方案,尽管分类任务不足以代表全部目标,但已证明该思路可行且具有前景。

原文摘要 · Abstract (English)

This position paper argues that the next generation of vision encoders should be image size agnostic and task driven. The source of our inspiration is biological. Not a structural aspect of biological vision, but a behavioral trait -- efficiency. We focus on a couple of ways in which vision in nature is efficient, but modern vision encoders not. We -- humans and animals -- deal with vast quantities of visual data, and need to be smart where we focus our limited energy -- it depends on the task. It is our belief that vision encoders should be dynamic and the computational complexity should depend on the task at hand rather than the size of the image. We, also, provide concrete first steps towards our vision -- a proof-of-concept solution for image classification. Despite classification being not very representative for what we are trying to achieve, it shows that our approach is feasible and promising.

视觉编码器任务驱动动态计算生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。