arXiv:2509.25805cs.CV2025-09

用动态图结构让SAM在极少数据下精准分割田间豌豆荚。

Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions

  • 构建可学习的动态相似性图,仅用400万参数实现高效适配
  • 在2到10样本下均显著提升分割精度,F-measure提升62.36%
  • 适合农业小目标检测场景,尤其适用于数据稀缺的田间应用

由于训练数据有限且田间环境复杂,基础模型在农业计算机视觉任务中的参数高效微调仍具挑战。本文提出基于动态相似性的图适应(DSGA)模块,用于在极端数据约束下适配分割一切模型(SAM),以实现复杂农业环境中小型密集目标的精确前景与实例分割。通过可学习多项式衰减初始化的权重排序机制构建动态相似性图,并结合自适应局部特征聚合,DSGA仅用400万可训练参数(仅为原始SAM的4.26%),建立了鲁棒的空间与动态相似性表征。将该图结构特征适配与低秩适配(LoRA)结合,形成互补优化框架,有效捕捉图像嵌入中的局部与全局依赖关系,同时保持模型稳定性和参数效率。在具有挑战性的豌豆荚数据集上的实验表明,DSGA+LoRA在2、4、8和10样本条件下均表现优异,性能随样本数增加持续提升。定量指标显示,结构度量提升17.31%,自适应F-measure提升62.36%。全面的消融实验与梯度加权类激活映射(Grad-CAM)、t-SNE可视化验证了该框架在特征区分上的有效性。所提方法在自动农业监测中具备实用价值,在田间条件下对含10至120个豆荚的图像实现了准确计数,调整后决定系数达0.8987。

原文摘要 · Abstract (English)

Parameter-Efficient Fine-Tuning (PEFT) of foundation models for agricultural computer vision tasks remains challenging due to limited training data and complex field conditions. This study introduces a Dynamic Similarity-based Graph Adaptation (DSGA) module to adapt the Segment Anything Model (SAM) under extreme data constraints for precise foreground and instance segmentation of small dense objects in complex agricultural environments. Through dynamic similarity graph construction with a learnable polynomial decay-initialized weight ranking mechanism and adaptive local feature aggregation, DSGA establishes robust spatial and dynamic similarity representation with only 4.00M trainable parameters, which is 4.26% of the original SAM. Integrating this graph-based feature adaptation with Low-Rank Adaptation (LoRA) creates a complementary optimization framework that effectively captures both local and global dependencies in image embeddings while preserving model stability and parameter efficiency. Experimental results on a challenging chickpea pod dataset demonstrated that DSGA with LoRA achieved superior performance across multiple metrics evaluated under 2, 4, 8 and 10 shots, with progressive performance gains as shot count increased. Quantitative metrics showed a 17.31% improvement in Structure-measure and a 62.36% gain in adaptive F-measure compared to the baseline SAM fine-tuning. Comprehensive ablation studies and visualization analyses through Grad-CAM and t-SNE validated the framework's effectiveness in feature discrimination. The proposed adaptation demonstrated practical utility for automated agricultural monitoring applications, achieving accurate pod-counting with an adjusted R-squared of 0.8987 for images with 10 to 120 pods under challenging field conditions.

小目标检测农业视觉参数高效动态图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。