为视觉大模型设计可拆解能力的评测基准,精准定位优劣。
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- 将14项基础视觉能力拆解独立评测,避免多能力混杂干扰
- 发现0.5B小模型与7B大模型在排名上相似,节省8倍算力
- 适合想精确评估模型视觉能力的研究者和工程师
视觉基础模型(VFMs)的发展亟需系统化评估。现有方法通常将VFMs与大语言模型(LLMs)结合,通过广泛VQA基准进行评估,但存在两大盲区:(i) 指令微调数据分布与VQA测试集不一致,错误可能源于数据偏差而非视觉能力不足;(ii) VQA任务常需多项视觉能力协同,难以判断失败是因缺全部能力还是单一关键项。为此,我们提出AVA-Bench,首个将14项原子视觉能力(AVAs)显式解耦的基准,如定位、深度估计、空间理解等,这些是复杂视觉推理的基础。通过在每项能力中匹配训练与测试分布,AVA-Bench可精准定位模型强弱。对主流VFMs的测评揭示了独特的“能力指纹”,使模型选择从经验猜测变为科学工程。值得注意的是,使用0.5B LLM即可获得与7B LLM相当的排序结果,算力消耗降低8倍,实现高效评估。通过提供全面透明的评测体系,我们期望AVA-Bench为下一代视觉基础模型奠定基础。
原文摘要 · Abstract (English)
The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as general-purpose heads, followed by evaluation on broad Visual Question Answering (VQA) benchmarks. However, this protocol has two key blind spots: (i) the instruction tuning data may not align with VQA test distributions, meaning a wrong prediction can stem from such data mismatch rather than a VFM' visual shortcomings; (ii) VQA benchmarks often require multiple visual abilities, making it hard to tell whether errors stem from lacking all required abilities or just a single critical one. To address these gaps, we introduce AVA-Bench, the first benchmark that explicitly disentangles 14 Atomic Visual Abilities (AVAs) -- foundational skills like localization, depth estimation, and spatial understanding that collectively support complex visual reasoning tasks. By decoupling AVAs and matching training and test distributions within each, AVA-Bench pinpoints exactly where a VFM excels or falters. Applying AVA-Bench to leading VFMs thus reveals distinctive "ability fingerprints," turning VFM selection from educated guesswork into principled engineering. Notably, we find that a 0.5B LLM yields similar VFM rankings as a 7B LLM while cutting GPU hours by 8x, enabling more efficient evaluation. By offering a comprehensive and transparent benchmark, we hope AVA-Bench lays the foundation for the next generation of VFMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。