不用微调就能预测遥感大模型性能,省时省力
Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
- 用能力编码方法捕捉模型核心能力,无需微调即可预测表现
- 在75个遥感大模型上验证,可准确预判其在多个下游任务的表现
- 适合遥感研究者快速选型模型,也帮助发现新研究方向
基础模型在计算机视觉中取得显著进展:经过一次昂贵的训练后,可应对多种任务。在地球观测领域,过去四年已开发超过75个遥感视觉基础模型。然而,尚无一个模型在所有下游任务中持续领先。为促进模型比较,我们提出一种低成本方法,可在无需对每个任务进行微调的前提下,预测模型在多个下游任务中的表现。该方法基于我们提出的“能力编码”(capabilities encoding)。该方法具有双重价值:一方面可简化新任务下基础模型的选择;另一方面用于重新审视现有文献,揭示未来研究方向。代码已开源:https://github.com/pierreadorni/capabilities-encoding。
原文摘要 · Abstract (English)
Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation models have been developed in the past four years. However, none has consistently outperformed the others across all available downstream tasks. To facilitate their comparison, we propose a cost-effective method for predicting a model's performance on multiple downstream tasks without the need for fine-tuning on each one. This method is based on what we call "capabilities encoding." The utility of this novel approach is twofold: we demonstrate its potential to simplify the selection of a foundation model for a given new task, and we employ it to offer a fresh perspective on the existing literature, suggesting avenues for future research. Codes are available at https://github.com/pierreadorni/capabilities-encoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。