用不完美的标注训练单目高度估计模型,提升跨区域泛化能力。
Enhancing Monocular Height Estimation via Weak Supervision from Imperfect Labels
- 基于弱监督设计集成框架,利用噪声标签中的有效信息
- 在DFC23和GBH数据集上平均RMSE降低超18%
- 适合缺乏高质量标注的遥感场景下的大范围高度估计
单目高度估计为遥感三维感知提供了高效低成本的解决方案。然而,训练深度神经网络需要大量标注数据,而高质量标签稀缺且仅在发达地区可用,限制了模型的泛化能力与大规模应用。本文通过利用域外区域的不完美标签来训练像素级高度估计网络,这些标签可能不完整、不精确或不准确。我们提出一种兼容任意单目高度估计网络的集成式流程,包含专为弱监督设计的架构与损失函数,通过平衡软损失和序数约束挖掘噪声标签中的信息。在两个数据集(DFC23,0.5–1 m;GBH,3 m)上的实验表明,该方法显著提升了跨域一致性,相较于基线,平均RMSE在DFC23上降低22.94%,在GBH上降低18.62%。消融实验验证了各设计组件的有效性。
原文摘要 · Abstract (English)
Monocular height estimation provides an efficient and cost-effective solution for three-dimensional perception in remote sensing. However, training deep neural networks for this task demands abundant annotated data, while high-quality labels are scarce and typically available only in developed regions, which limits model generalization and constrains their applicability at large scales. This work addresses the problem by leveraging imperfect labels from out-of-domain regions to train pixel-wise height estimation networks, which may be incomplete, inexact, or inaccurate compared to high-quality annotations. We introduce an ensemble-based pipeline compatible with any monocular height estimation network, featuring architecture and loss functions specifically designed to leverage information in noisy labels through weak supervision, utilizing balanced soft losses and ordinal constraints. Experiments on two datasets -- DFC23 (0.5--1 m) and GBH (3 m) -- show that our method achieves more consistent cross-domain performance, reducing average RMSE by up to 22.94% on DFC23 and 18.62% on GBH compared with baselines. Ablation studies confirm the contribution of each design component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。