arXiv:2507.10571cs.AIcs.CL2025-07被引 7

用可信调度与检索增强推理,让AI在零样本下也能准确识别苹果叶病。

Agentic AI with Orchestrator-Agent Trust: A Modular Visual Classification Framework with Trust-Aware Orchestration and RAG-Based Reasoning

  • 分模块设计:视觉代理负责感知,调度器做元推理,结合RAG提升可信度。
  • 零样本下准确率提升77.94%,达85.63%,优于调优模型表现。
  • 适合医疗、生物等需高可信度的场景,支持可解释与可扩展系统构建。

现代人工智能越来越多依赖融合视觉与语言理解的多智能体架构。然而核心挑战在于:在无微调的零样本设置下如何建立对智能体的信任?本文提出一种新型模块化智能体视觉分类框架,整合通用多模态智能体、非视觉推理调度器与基于检索增强生成(RAG)的模块。以苹果叶病诊断为例,对比三种配置:(I) 基于置信度调度的零样本方法;(II) 微调后智能体;(III) 结合CLIP图像检索与重评估循环的可信校准调度。通过ECE、OCR、CCC等置信度校准指标,调度器动态调节各智能体信任度。结果表明,采用可信调度与RAG后,零样本设置下准确率提升77.94%,整体达到85.63%。GPT-4o表现更佳校准性,而Qwen-2.5-VL存在过度自信。图像-RAG通过引入视觉相似案例,实现预测修正与迭代重评。该系统分离感知与元推理,具备可扩展性与可解释性,为诊断、生物学等高可信需求领域提供可靠智能体范式。

原文摘要 · Abstract (English)

Modern Artificial Intelligence (AI) increasingly relies on multi-agent architectures that blend visual and language understanding. Yet, a pressing challenge remains: How can we trust these agents especially in zero-shot settings with no fine-tuning? We introduce a novel modular Agentic AI visual classification framework that integrates generalist multimodal agents with a non-visual reasoning orchestrator and a Retrieval-Augmented Generation (RAG) module. Applied to apple leaf disease diagnosis, we benchmark three configurations: (I) zero-shot with confidence-based orchestration, (II) fine-tuned agents with improved performance, and (III) trust-calibrated orchestration enhanced by CLIP-based image retrieval and re-evaluation loops. Using confidence calibration metrics (ECE, OCR, CCC), the orchestrator modulates trust across agents. Our results demonstrate a 77.94\% accuracy improvement in the zero-shot setting using trust-aware orchestration and RAG, achieving 85.63\% overall. GPT-4o showed better calibration, while Qwen-2.5-VL displayed overconfidence. Furthermore, image-RAG grounded predictions with visually similar cases, enabling correction of agent overconfidence via iterative re-evaluation. The proposed system separates perception (vision agents) from meta-reasoning (orchestrator), enabling scalable and interpretable multi-agent AI. This blueprint illustrates how Agentic AI can deliver trustworthy, modular, and transparent reasoning, and is extensible to diagnostics, biology, and other trust-critical domains. In doing so, we highlight Agentic AI not just as an architecture but as a paradigm for building reliable multi-agent intelligence. agentic ai, orchestrator agent trust, trust orchestration, visual classification, retrieval augmented reasoning

智能体可信推理视觉分类RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。