多智能体协作解决农业多图问答难题,支持动态迭代与跨图信息融合。
Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
- 四角色智能体协同:检索、反思、答题、优化,实现动态推理
- 在AgMMU上表现优于基线,支持多图互补与上下文增强
- 适合需要多源图像分析的农业智能问答场景
农业视觉问答对农民和研究人员提供及时准确的知识至关重要。然而,现有方法多针对文本仅限或单图场景设计,难以应对真实农业中需多图输入、跨空间尺度与生长阶段互补视图的复杂需求。同时,受限于实时外部农业知识获取,系统在证据不全时难以适应。此外,固定流程缺乏系统性质量控制。为此,我们提出一个自省与自提升的多智能体框架,包含检索者、反思者、答题者与改进者四个角色。检索者生成查询并获取外部信息,反思者评估充分性并触发重构与再检索;两名答题者并行生成候选答案以降低偏见;改进者通过迭代校验优化答案,并确保多图信息有效对齐与利用。在AgMMU基准测试中,该框架展现出竞争力的多图像农业问答性能。
原文摘要 · Abstract (English)
Agricultural visual question answering is essential for providing farmers and researchers with accurate and timely knowledge. However, many existing approaches are predominantly developed for evidence-constrained settings such as text-only queries or single-image cases. This design prevents them from coping with real-world agricultural scenarios that often require multi-image inputs with complementary views across spatial scales, and growth stages. Moreover, limited access to up-to-date external agricultural context makes these systems struggle to adapt when evidence is incomplete. In addition, rigid pipelines often lack systematic quality control. To address this gap, we propose a self-reflective and self-improving multi-agent framework that integrates four roles, the Retriever, the Reflector, the Answerer, and the Improver. They collaborate to enable context enrichment, reflective reasoning, answer drafting, and iterative improvement. A Retriever formulates queries and gathers external information, while a Reflector assesses adequacy and triggers sequential reformulation and renewed retrieval. Two Answerers draft candidate responses in parallel to reduce bias. The Improver refines them through iterative checks while ensuring that information from multiple images is effectively aligned and utilized. Experiments on the AgMMU benchmark show that our framework achieves competitive performance on multi-image agricultural QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。