对比通用与专用视觉模型,发现专用模型会损失泛化能力。
How Universal Are SAM2 Features?

- 用可训练轻量颈部测试冻结特征的适应性
- 专用模型在深度估计上表现更好,但泛化任务更差
- 逐层分析揭示表征瓶颈,适合模型设计参考
通用基础视觉模型与专用模型之间的权衡对高效特征编码设计至关重要,但尚未完全理解。本文通过比较通用型Hiera编码器与分割专用型Segment Anything Model 2(SAM2)的特征通用性,利用轻量可训练颈部探测其冻结特征的适应能力,量化了专业化带来的信息论代价。结果表明,尽管SAM2在深度估计等空间相关任务中表现优异,但其在姿态估计、图像描述等概念相距较远的任务上表现逊于通用型的Hiera,显示出更广泛的语义信息损失。对SAM2的新型跨颈部分析进一步揭示:每层级的适应都会引入新的表征瓶颈。该研究为多样化下游应用的高效特征编码与适配策略提供了定量依据。
原文摘要 · Abstract (English)
The trade-off between general-purpose foundation vision models and their specialized counterparts is critical for efficient feature coding design and is not yet fully understood. We investigate this trade-off by comparing the feature versatility of the general-purpose Hiera encoder against the segmentation-specialized Segment Anything Model 2 (SAM2). Using a lightweight, trainable neck to probe the adaptability of their frozen features, we quantify the information-theoretic cost of specialization. Our results reveal that while SAM2's specialization is highly effective for spatially-related tasks like depth estimation, it comes at a cost. The specialized SAM2 encoder underperforms its generalist predecessor, Hiera, on conceptually distant tasks such as pose estimation and image captioning, demonstrating a measurable loss of broader semantic information. A novel cross-neck analysis on SAM2 reveals that each level of adaptation creates a further representational bottleneck. Our analysis illuminates these trade-offs in feature universality, providing a quantitative foundation for designing efficient feature coding and adaptation strategies for diverse downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。