arXiv:2605.08213cs.CV2026-05

用低成本双目相机实现无人机对细枝的精准三维定位,无需额外传感器。

Low-Cost Stereo Vision for Robust 3D Positioning of Thin Radiata Pine Branches in Autonomous Drone Pruning

论文配图:Low-Cost Stereo Vision for Robust 3D Positioning of Thin Radiata Pine Branches in Autonomous Drone Pruning
图 1 · 摘自论文原文
  • 用双目视觉+分割算法识别树冠中10毫米粗的细枝
  • 通过三角测量与异常值剔除,提升薄枝深度估计精度
  • 适合林业自动化修剪场景,降低设备成本

新西兰主要经济树种辐射松的人工修剪危险且耗人力,当前自主修剪平台多依赖昂贵的激光雷达,且仅适用于粗枝,难以推广。本文研究在无人机上仅用单个低成本双目相机(ZED Mini)能否实现对直径10毫米细枝的精确检测与三维定位,从而替代辅助深度传感器。提出两阶段流程:分支分割与深度估计。分割采用自建71对立体图像数据集,对比Mask R-CNN、YOLOv8和YOLOv9;选择YOLOv8/v9作为实时分割代表,并设计可兼容后续YOLO版本。深度估计对比传统方法SGBM+WLS与六种深度学习模型(PSMNet、ACVNet、GWCNet、MobileStereoNet、RAFT-Stereo、NeRF-Supervised Deep Stereo),并进行跨数据集微调实验,揭示城市驾驶基准与自然林地场景间的域差距。创新点在于将分割掩码与视差图结合,采用质心三角化与中位数绝对偏差异常值剔除法,生成鲁棒的枝条到相机距离,有效应对森林场景中纹理稀疏、结构纤细、视差噪声等问题。1-2米距离下定性评估显示,基于学习的立体方法产生更一致的深度估计。

原文摘要 · Abstract (English)

Manual pruning of radiata pine, a species of major economic importance to New Zealand forestry, is hazardous, labour-intensive, and increasingly constrained by workforce shortages. Existing autonomous pruning platforms typically rely on expensive sensors such as LiDAR and are limited to thick branches, which restricts their wider adoption. This paper investigates whether a single low-cost stereo camera mounted on a drone can provide sufficiently accurate branch detection and three-dimensional positioning to support autonomous pruning of branches as thin as 10 mm, thereby removing the need for auxiliary depth sensors. The proposed pipeline comprises two stages: branch segmentation and depth estimation. For segmentation, Mask R-CNN variants and the YOLOv8 and YOLOv9 families are compared on a custom dataset of 71 stereo image pairs captured with a ZED Mini camera; YOLOv8 and YOLOv9 are selected as representative state-of-the-art real-time segmentors at the time of data collection, and the framework is designed to remain compatible with newer YOLO releases. For depth estimation, a traditional method (SGBM with WLS filtering) and deep-learning-based methods (PSMNet, ACVNet, GWCNet, MobileStereoNet, RAFT-Stereo, and NeRF-Supervised Deep Stereo) are evaluated, including cross-dataset fine-tuning experiments that expose the domain gap between urban driving benchmarks and natural forestry scenes. The main novelty of this work lies in coupling stereo segmentation with a centroid-based triangulation algorithm and Median-Absolute-Deviation outlier rejection that converts a segmentation mask and disparity map into a single robust branch-to-camera distance, addressing the challenges of sparse texture, thin structures, and noisy disparity values typical of forest scenes. Qualitative evaluations at distances of 1-2 m show that the learning-based stereo methods produce more coherent depth es...

无人机立体视觉林业机器人细枝检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。