arXiv:2609.06051cs.CVcs.AI2026-09

发现图像识别模型的缩放鲁棒性主要由基础准确率决定,而非模型大小或架构。

Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis

论文配图:Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis
图 1 · 摘自论文原文
  • 通过系统降尺度测试20个主流模型,定义特征缩放点衡量识别性能下降起点。
  • 基础准确率越高,特征缩放点越低(相关系数-0.89),且该规律稳定可靠。
  • 模型参数量和架构家族对缩放鲁棒性无显著影响,适合关注模型泛化能力的研究者。

视觉识别模型对图像缩放的敏感性已广为人知,但不同架构间该敏感性的成因仍不明确。本文评估了20个预训练的ImageNet-1K分类器,涵盖7类主流架构(卷积、移动、高效、Transformer等)。通过系统降低输入图像尺度,构建缩放-准确率响应曲线,并定义特征缩放点作为识别性能显著下降的紧凑度量。结果表明,基础准确率与特征缩放点呈强负相关(皮尔逊相关系数r = -0.890,R² = 0.792,p < 10⁻⁶),该关系在自助抽样、留一架构和留一族分析中均保持稳定。相比之下,控制基础准确率后,参数量无额外解释力(p = 0.80);架构族亦无显著增量解释力。此外,特征缩放点与表征稳定性几乎无关(r = -0.003,p = 0.991)。结果表明,所研究模型中,缩放鲁棒性主要由基础识别性能决定,而非模型规模、架构类型或表征稳定性。本研究提供了跨架构缩放鲁棒性的实证框架,并揭示了一个可复现的精度-缩放规律,值得进一步理论探索。

原文摘要 · Abstract (English)

The sensitivity of visual recognition models to changes in image scale is well established, yet the factors governing this sensitivity across heterogeneous architectures remain unclear. In this work, we investigate whether scale robustness exhibits a common quantitative structure across modern vision models. We evaluate 20 pretrained ImageNet-1K classifiers spanning seven architectural families, including convolutional, mobile, efficient, and Transformer-based architectures. By systematically reducing input image scale, we construct scale-accuracy response curves and define a characteristic scale as a compact measure of the onset of substantial recognition degradation. We then examine the relationship between characteristic scale and baseline recognition accuracy, model parameter count, architectural family, and representation stability. A strong inverse association is observed between baseline accuracy and characteristic scale (Pearson r = -0.890, R^2= 0.792, p < 10^-6). This relationship remains stable under bootstrap resampling, leave-one-architecture-out analysis, and leave-one-family-out analysis. In contrast, parameter count provides negligible additional explanatory power after controlling for baseline accuracy (p = 0.80), while architectural family does not provide significant incremental explanatory power. Furthermore, characteristic scale shows essentially no association with representation stability (r = -0.003, p = 0.991). These results indicate that, across the studied models, scale robustness is strongly organized by baseline recognition performance rather than simply by model size, architectural family, or representation stability. The study provides an empirical framework for characterizing scale robustness across vision architectures and identifies a reproducible accuracy-scale regularity that warrants further theoretical investigation.

视觉识别缩放鲁棒性模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。