arXiv:2605.20892cs.CV2026-05

用多模型协作+大模型纠错,提升水果细粒度识别准确率

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

论文配图:FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
图 1 · 摘自论文原文
  • 两阶段动态推理:先集成多个模型,再用大模型处理不确定样本
  • 在306类11.6万张图像上达到70.49%准确率,优于现有方法
  • 适合农业视觉分拣、品质检测等实际部署场景

细粒度水果分类是农业计算机视觉中的关键挑战,主要受限于高质量数据集稀缺及类别间视觉差异小。为此,我们构建了包含306个水果类别、共116,233张样本的综合性数据集。提出FruitEnsemble,一种面向异构模型集成的实用两阶段动态推理框架,以克服静态单模型架构的泛化局限。第一阶段通过验证校准的加权集成生成鲁棒的Top-3候选池;针对困难样本,引入专家仲裁机制:当集成置信度低于0.6时,触发多模态大语言模型(MLLM),结合外部植物学描述,通过思维链(CoT)推理进行严格视觉验证。此外,采用硬样本感知联合损失优化训练流程。大量实验表明,FruitEnsemble达到70.49%分类准确率,优于现有最先进模型。该框架为真实世界农业视觉分拣与品质检测任务提供高效、可部署的解决方案。

原文摘要 · Abstract (English)

Fine-grained fruit classification is a critical yet challenging task in agricultural computer vision, primarily hindered by a severe shortage of high-quality datasets and the high visual similarity between classes. To address these challenges, we first constructed a comprehensive dataset comprising 306 fruit categories with 116,233 samples. Moreover, we propose FruitEnsemble, a practical two-stage dynamic inference framework designed to overcome the generalization limitations of static single-model architectures. In the first stage, FruitEnsemble employs a validation-calibrated weighted ensemble of heterogeneous backbones to generate a robust Top-3 candidate pool. To tackle difficult samples, we introduce an expert arbitration mechanism: when ensemble confidence falls below 0.6, a multimodal large language model (MLLM) is triggered to perform rigorous visual verification by integrating external botanical descriptions using Chain-of-Thought (CoT) reasoning. Furthermore, we optimized the training pipeline with a hard sample-aware joint loss. Extensive experiments demonstrate that FruitEnsemble achieves a classification accuracy of 70.49\% and outperforms existing state-of-the-art models. Our framework provides an efficient, deployment-oriented solution for real-world agricultural visual sorting and quality inspection tasks.

细粒度识别多模型集成大模型应用农业视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。