无需标注数据,实时检测农田机器人前方隐藏障碍物。
Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover

- 用记忆增强的视频变换器捕捉动态场景中的正常视觉模式
- 在油菜田数据集上达0.973检测与0.997分割AUC,性能领先
- 轻量版推理仅需14毫秒,适合机器人实时安全决策
尽管自主农业机器人在精准农业中日益重要,但保障持续运行安全仍是关键挑战。传统传感器如激光雷达无法探测植株冠层下方的障碍物,构成重大风险。基于摄像头的监督学习方法虽能识别常见物体,但在训练数据中未出现的新障碍物上表现不佳。无监督异常检测通过学习环境正常视觉模式提供解决方案,但对移动机器人拍摄的动态场景常失效。本文提出视频记忆变换器异常检测(VMTAD),一种完全无监督的实时障碍物检测方法。该方法采用基于变换器的架构并引入专用记忆模块,利用前序帧编码表示处理时间上下文,有效应对机器人运动带来的动态变化。模型仅使用正常运行图像进行训练,无需标签。在'Grillion'农业机器人上评估显示,在具有挑战性的油菜数据集上,VMTAD达到0.973检测和0.997分割的接收者操作特征曲线下面积(AUC)。其轻量版本实现高精度与实时推理(14毫秒)的平衡,经分析确认足以满足机器人总制动距离的安全需求。
原文摘要 · Abstract (English)
While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challenge. Conventional safety sensors, such as LiDAR, fail to detect obstacles positioned below the plant canopy, posing a significant risk. While camera-based supervised learning methods can detect common objects, they perform poorly when faced with obstacles that were not present in their training data. Actual unsupervised anomaly detection offers a solution by learning the normal visual patterns of an environment, but often fails for the dynamic scenes captured by a moving rover.\\ This paper introduces Video Memory Transformers for Anomaly Detection (VMTAD), a fully unsupervised method designed for real-time obstacle detection in dynamic agricultural scenes. VMTAD utilizes a transformer-driven architecture augmented with a dedicated memory module. This memory module leverages temporal context by processing encoded representations of preceding frames. This approach enables the system to effectively address the dynamic context caused by the robot's movement. The model is trained using only images that represent normal operation, requiring no data labels.\\ VMTAD was rigorously evaluated on the 'Grillion' agricultural rover. On a challenging rapeseed dataset, VMTAD achieved state-of-the-art performance, reaching a 0.973 detection and 0.997 segmentation Area Under the Receiver Operating Characteristic curve. A lightweight variant provides an optimal balance of high accuracy and real-time inference (14 ms), which is critical for safety, as confirmed by our analysis of the rover's total stopping distance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。