让视觉模型像专家一样推理,通过知识增强实现开放场景下的细粒度识别。
Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
- 构建三阶段闭环推理框架,结合网络检索与视觉证据定位。
- 在开放集下推理准确率提升19%,显著优于现有模型。
- 适合需要可解释性推理的细粒度视觉任务,如生物分类、医学影像分析。
细粒度视觉理解正从静态分类转向知识增强的推理,要求模型不仅能识别,还能提供依据。现有方法受限于封闭类别体系和单标签预测,在开放集或上下文依赖条件下性能显著下降。本文提出知识增强的细粒度推理代理(KFRA),将细粒度感知转化为证据驱动的推理。KFRA采用三阶段闭合推理循环,模拟专家分析流程:首先进行开集检测与网络规模检索生成类别假设;其次通过全局到局部聚焦机制,将文本知识与视觉证据对齐,定位判别区域;最后在大型多模态模型中整合多模态证据进行可解释推理。不同于将检索与推理分离的现有代理,KFRA建立检索-接地耦合机制,将检索知识转化为空间对齐的验证证据。该设计实现跨任务、跨场景的事实性、可解释性推理。为评估能力,构建了包含六个知识维度的FGExpertBench基准。大量实验表明,KFRA在推理准确率上相较独立的大规模多模态模型及当前代理框架平均提升19%,在开放集细粒度视觉理解中实现证据锚定的可解释性。
原文摘要 · Abstract (English)
Fine-grained visual understanding is shifting from static classification to knowledge-augmented reasoning, where models must justify as well as recognise. Existing approaches remain limited by closed-set taxonomies and single-label prediction, leading to significant degradation under open-set or context-dependent conditions. We present the Knowledge-Augmented Fine-Grained Reasoning Agent (KFRA), a unified framework that transforms fine-grained perception into evidence-driven reasoning. KFRA operates through a three-stage closed reasoning loop that emulates expert analysis. It first performs open-vocabulary detection and web-scale retrieval to generate category hypotheses. It then conducts discriminative regions localisation by aligning textual knowledge with visual evidence through a global-to-local focusing mechanism. Finally, it integrates all multimodal evidence within a large multimodal model to perform interpretable reasoning. Unlike existing agents that treat retrieval and reasoning as independent processes, KFRA establishes a retrieval-grounding coupling that converts retrieved knowledge into spatially grounded evidence for verification. This design enables factual, interpretable, and task-agnostic reasoning across diverse fine-grained scenarios. To evaluate this capability, we construct FGExpertBench, a benchmark designed to assess reasoning depth and cross-task generalisation across six knowledge dimensions. Extensive experiments demonstrate that KFRA consistently surpasses both standalone large multimodal models and current agent frameworks, achieving up to 19 percent improvement in reasoning accuracy and delivering evidence-grounded interpretability in open-set fine-grained visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。