arXiv:2504.13365cs.CVcs.AI2025-04被引 18

用视觉语言模型+联邦学习,让农场数据不外泄还能高效训练检测模型

VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture

  • 用轻量级提示生成器提升视觉语言模型性能,适配分布式农场数据
  • 相比传统方法,检测精度提升14.53%,通信开销降低99.3%
  • 适合隐私敏感的智慧农业场景,支持多类农情目标识别

在现代智慧农业中,目标检测对实现自动化、精准种植和资源监控至关重要。从识别作物健康状况与虫害,到优化收获流程,精准的目标检测可显著提升生产效率与可持续性。然而,训练目标检测模型通常需要大规模数据收集,且在各农场分布的敏感农业数据下存在隐私风险。为此,我们提出VLLFL——一种基于视觉语言模型的轻量化联邦学习框架。该框架利用视觉语言模型的泛化能力和上下文感知检测能力,并结合联邦学习的隐私保护特性。通过训练一个紧凑的提示生成器,增强部署在不同农场的视觉语言模型性能,既保障隐私又大幅降低通信开销。实验表明,VLLFL在视觉语言模型性能上提升14.53%,同时通信开销减少99.3%。该框架可覆盖从多种水果识别到农业有害动物检测等任务,为农业应用提供高效、可扩展且隐私安全的解决方案。

原文摘要 · Abstract (English)

In modern smart agriculture, object detection plays a crucial role by enabling automation, precision farming, and monitoring of resources. From identifying crop health and pest infestations to optimizing harvesting processes, accurate object detection enhances both productivity and sustainability. However, training object detection models often requires large-scale data collection and raises privacy concerns, particularly when sensitive agricultural data is distributed across farms. To address these challenges, we propose VLLFL, a vision-language model-based lightweight federated learning framework (VLLFL). It harnesses the generalization and context-aware detection capabilities of the vision-language model (VLM) and leverages the privacy-preserving nature of federated learning. By training a compact prompt generator to boost the performance of the VLM deployed across different farms, VLLFL preserves privacy while reducing communication overhead. Experimental results demonstrate that VLLFL achieves 14.53% improvement in the performance of VLM while reducing 99.3% communication overhead. Spanning tasks from identifying a wide variety of fruits to detecting harmful animals in agriculture, the proposed framework offers an efficient, scalable, and privacy-preserving solution specifically tailored to agricultural applications.

智慧农业联邦学习视觉语言模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。