构建新基准,评估模型主动获取细粒度知识的能力
FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition

- 提出可验证证据的细粒度知识获取任务框架
- 顶尖模型准确率仅25.1%,超30%者无一
- 失败主因是实体检索错误与视觉判断差
日常中的细粒度识别常非闭卷分类:面对陌生物体时,人类会主动搜索、比对细节并验证证据。现有基准多聚焦视觉识别,忽视主动获取外部知识的能力。本文研究细粒度知识获取,要求系统通过搜寻、验证和利用外部证据回答开放性细粒度识别问题。我们提出FIKA-Bench,一个包含311个公开来源与真实场景样本的泄漏感知、证据支撑数据集。每例均经前沿闭卷模型筛选以去除记忆样本,并人工审核排除图像-答案泄露,确保仅保留有可靠证据支持的样本。对最新大型多模态模型(LMMs)及智能体的评估显示,该任务仍具极大挑战:最佳系统准确率仅25.1%,无模型突破30%。关键发现:仅赋予工具不足以缩小差距;智能体失败主要源于错误实体检索与差劲视觉判断。结果表明,可靠知识获取需更注重细粒度识别的智能体设计。
原文摘要 · Abstract (English)
Fine-grained recognition in everyday life is often not a closed-book classification problem: when encountering unfamiliar objects, humans actively search, compare visual details, and verify evidence before deciding. Existing benchmarks primarily evaluate visually recognition, leaving this active external knowledge acquisition ability underexplored. We study fine-grained knowledge acquisition, where a system must seek, verify, and use external evidence to answer open-ended fine-grained recognition questions. We introduce FIKA-Bench, a leakage-aware and evidence-grounded collection of 311 public-source and real-life instances. To ensure high quality, every example is filtered against frontier closed-book models to remove memorized cases and audited to eliminate image-answer leakage, retaining only samples supported by verified evidence. Our evaluation of latest Large Multimodal Models (LMMs) and agents reveals that the task remains a formidable challenge: the best system reaches only 25.1% accuracy, with no model exceeding 30%. Crucially, we find that merely equipping models with tools is insufficient to bridge this gap; agent failures are predominantly driven by wrong entity retrieval and poor visual judgement. These results show that reliable knowledge acquisition needs better agent designs that focus on fine-grained recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。