arXiv:2608.24935cs.CVcs.AI2026-08

轻量级多模态模型精准识别苹果幼果结构,助力果园机器人早期疏果。

A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

论文配图:A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
图 1 · 摘自论文原文
  • 基于TinyCLIP改进,用领域提示词实现图像与果树结构对齐
  • 三类结构分类F1最高达0.98,整体宏平均F1达0.93
  • 可部署于Jetson设备,模型仅127-137MB,推理毫秒级

准确识别苹果幼果的花萼、果体和果柄等解剖结构,对机器人疏果、产量管理等精准果园作业至关重要。本研究提出一种轻量级多模态视觉语言框架,适配TinyCLIP以实现复杂果园环境下细粒度幼果结构分类。从Scilate和Scifresh果园采集600张高分辨率RGB图像,裁剪为224×224图像块并标注三类解剖结构。采用领域特定语言提示(如“一张某类的照片”)引导图像与园艺结构的多模态对齐。通过步长为112像素的滑动窗口推理策略,将块级预测聚合为空间热图,实现与机器人疏果相关的幼果组件可解释定位。在NVIDIA T4 GPU上块级评估,花萼F1为0.95,果体为0.98,果柄为0.85,宏平均F1为0.93。采用ONNX与TensorRT进行部署优化,在NVIDIA Jetson硬件上实现高效推理,支持约127–137 MB模型大小,并在INT8量化下保持精度,单块推理达毫秒级。结果表明,轻量级视觉语言模型可为自动化幼果分析及未来机器人疏果系统提供可解释且边缘部署的感知能力。源码与实现细节已公开于https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards。

原文摘要 · Abstract (English)

Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.

视觉语言果实识别边缘计算机器人农业

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。