首次评估深度立体匹配方法在林业无人机场景下的泛化能力。
Generalization Evaluation of Deep Stereo Matching Methods for UAV-Based Forestry Applications
- 在无训练情况下测试8种先进立体匹配模型,覆盖迭代优化与基础模型范式。
- 林业数据集上DEFOM表现最佳,结构化场景中误差低至0.35-4.65像素。
- 部分模型在特定场景崩溃(如RAFT-Stereo在ETH3D误差高达26.23像素),警示跨域风险。
自主无人机林木作业需要具备强跨域泛化能力的深度估计方法,但现有评估多集中于城市和室内场景,缺乏对植被密集环境的系统评价。本文首次对八种前沿立体匹配方法(RAFT-Stereo、IGEV、IGEV++、BridgeDepth、StereoAnywhere、DEFOM及基线方法ACVNet、PSMNet、TCstereo)进行零样本系统评估,这些方法涵盖迭代优化、基础模型与零样本适应范式。所有模型仅在Scene Flow数据集上训练,未做微调,测试于四个标准基准(ETH3D、KITTI 2012/2015、Middlebury)以及新构建的5,313对坎特伯雷林业数据集(使用ZED Mini相机采集,分辨率1920x1080)。结果揭示场景依赖性:基础模型在结构化场景中表现优异(BridgeDepth:ETH3D为0.23像素,KITTI为0.83-1.07像素;DEFOM:0.35-4.65像素),而迭代方法更具跨域鲁棒性(IGEV++:0.36-6.77像素;IGEV:0.33-21.91像素)。关键发现:RAFT-Stereo在ETH3D出现灾难性失败(平均端点误差26.23像素,错误率98%),源于负视差预测;在KITTI上表现正常(0.90-1.11像素)。坎特伯雷林业数据集的定性分析表明,尽管IGEV++细节更精细,但DEFOM在深度平滑性、遮挡处理和跨域一致性方面更优,是植被深度估计的理想基线。
原文摘要 · Abstract (English)
Autonomous UAV forestry operations require robust depth estimation methods with strong cross-domain generalization. However, existing evaluations focus on urban and indoor scenarios, leaving a critical gap for specialized vegetation-dense environments. We present the first systematic zero-shot evaluation of eight state-of-the-art stereo methods--RAFT-Stereo, IGEV, IGEV++, BridgeDepth, StereoAnywhere, DEFOM (plus baseline methods ACVNet, PSMNet, TCstereo)--spanning iterative refinement, foundation model, and zero-shot adaptation paradigms. All methods are trained exclusively on Scene Flow and evaluated without fine-tuning on four standard benchmarks (ETH3D, KITTI 2012/2015, Middlebury) plus a novel 5,313-pair Canterbury forestry dataset captured with ZED Mini camera (1920x1080). Performance reveals scene-dependent patterns: foundation models excel on structured scenes (BridgeDepth: 0.23 px on ETH3D, 0.83-1.07 px on KITTI; DEFOM: 0.35-4.65 px across benchmarks), while iterative methods maintain cross-domain robustness (IGEV++: 0.36-6.77 px; IGEV: 0.33-21.91 px). Critical finding: RAFT-Stereo exhibits catastrophic ETH3D failure (26.23 px EPE, 98 percent error rate) due to negative disparity predictions, while performing normally on KITTI (0.90-1.11 px). Qualitative evaluation on Canterbury forestry dataset identifies DEFOM as the optimal gold-standard baseline for vegetation depth estimation, exhibiting superior depth smoothness, occlusion handling, and cross-domain consistency compared to IGEV++, despite IGEV++'s finer detail preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。