arXiv:2601.19461cs.CVcs.RO2026-01被引 3

首个针对树枝场景的深度估计基准测试,验证DEFOM为最佳基线。

Towards Gold-Standard Depth Estimation for Tree Branches in UAV Forestry: Benchmarking Deep Stereo Matching Methods

  • 系统评估8种立体匹配方法在树丛环境中的表现,使用预训练权重零样本测试。
  • DEFOM在树枝数据集上一致领先,平均排名1.75,跨域一致性最强。
  • 首次构建5313对高分辨率树枝深度数据集,填补植被密集场景评估空白。

自主无人机林业作业需要具备强跨域泛化能力的深度估计,但现有评估集中于城市和室内场景,缺乏对植被密集环境的覆盖。本文首次系统性地进行零样本评估,涵盖八种立体匹配方法,包括迭代优化、基础模型、扩散模型和3D CNN范式。所有方法均使用官方发布的预训练权重(在Scene Flow上训练),并在四个标准基准(ETH3D、KITTI 2012/2015、Middlebury)以及一个全新的5,313对坎特伯雷树枝数据集(1920×1080)上进行测试。结果揭示场景依赖模式:基础模型在结构化场景中表现优异(BridgeDepth: ETH3D上0.23像素;DEFOM: Middlebury上4.65像素),而迭代方法跨基准性能波动大(IGEV++: ETH3D上0.36像素,但Middlebury上达6.77像素;IGEV: ETH3D上0.33像素,但Middlebury上4.99像素)。在树枝数据集上的定性评估表明,DEFOM是植被深度估计的黄金标准基线,具有出色的跨域一致性(在各基准中始终排名第1-2,平均排名1.75)。其预测结果将作为未来基准测试的伪真值。

原文摘要 · Abstract (English)

Autonomous UAV forestry operations require robust depth estimation with strong cross-domain generalization, yet existing evaluations focus on urban and indoor scenarios, leaving a critical gap for vegetation-dense environments. We present the first systematic zero-shot evaluation of eight stereo methods spanning iterative refinement, foundation model, diffusion-based, and 3D CNN paradigms. All methods use officially released pretrained weights (trained on Scene Flow) and are evaluated on four standard benchmarks (ETH3D, KITTI 2012/2015, Middlebury) plus a novel 5,313-pair Canterbury Tree Branches dataset ($1920 \times 1080$). Results reveal scene-dependent patterns: foundation models excel on structured scenes (BridgeDepth: 0.23 px on ETH3D; DEFOM: 4.65 px on Middlebury), while iterative methods show variable cross-benchmark performance (IGEV++: 0.36 px on ETH3D but 6.77 px on Middlebury; IGEV: 0.33 px on ETH3D but 4.99 px on Middlebury). Qualitative evaluation on the Tree Branches dataset establishes DEFOM as the gold-standard baseline for vegetation depth estimation, with superior cross-domain consistency (consistently ranking 1st-2nd across benchmarks, average rank 1.75). DEFOM predictions will serve as pseudo-ground-truth for future benchmarking.

深度估计无人机林业立体匹配植被场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。