arXiv:2605.00908cs.CV2026-05被引 1

对比卷积与注意力模型在番茄田杂草检测中的表现

Evaluation of Convolutional and Transformer-Based Detectors for Weed Detection in Tomato Plantations

论文配图:Evaluation of Convolutional and Transformer-Based Detectors for Weed Detection in Tomato Plantations
图 1 · 摘自论文原文
  • 用YOLOv26-nano和RT-DETR等模型对比检测效果
  • 卷积模型更快更省资源,注意力模型更准但耗算力
  • 适合农业自动化中对速度与精度有不同需求的场景

本文对比评估了卷积神经网络与基于Transformer的目标检测架构在番茄种植园早期杂草检测中的表现。选用YOLOv26-nano(YOLO系列新变体)及RT-DETR Large、RF-DETR Medium等典型模型,在GROUNDBASED_WEED数据集上进行测试,涵盖六类杂草与未识别植物类别。通过精确率、召回率、平均精度及推理速度等指标,并结合非参数统计检验,结果表明:卷积模型在较低计算成本下实现较高性能,而基于Transformer的方法虽能更好捕捉全局上下文,但资源消耗更高。该研究为精准农业中的模型选型提供了实用依据。

原文摘要 · Abstract (English)

This paper presents a comparative evaluation of convolutional and transformer-based object detection architectures for early weed detection in tomato plantations. Representative models from each paradigm are considered, including YOLOv26-nano, a recent variant of the YOLO family, and RT-DETR Large and RF-DETR Medium as transformer-based architectures. The evaluation was conducted on the GROUNDBASED_WEED dataset, considering six weed classes and an additional category corresponding to unidentified plants, which allowed for the assessment of performance in terms of detection accuracy and computational efficiency using metrics such as precision, recall, average precision, and inference speed, as well as non-parametric statistical tests. The results highlight a clear trade-off between efficiency and contextual modeling: CNN-based detectors achieve high performance at a lower computational cost, while transformer-based approaches offer better global context capture at the expense of higher resource demands. These results provide practical criteria for model selection in precision agriculture applications.

目标检测农业智能视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。