用元初始化提升单图室内深度估计的跨数据集泛化能力。
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
- 将每个图像批次视为任务,用梯度元学习构建更强先验
- 在有限数据下使RMSE最高降低27.8%,跨数据集推理性能更优
- 可作为插件适配现有方法,适合部署于真实机器人场景
室内机器人依赖深度信息完成导航或避障等任务,单图深度估计广泛用于感知辅助。当前多数研究关注真实环境下的鲁棒性,但对未见数据集的泛化能力关注不足。本文采用基于梯度的元学习,在零样本跨数据集推理中实现更高泛化性。不同于带有明确类别标签的图像分类元学习,连续深度值与高度变化的室内环境(物体布局、场景构成)关联,无显式任务边界。我们提出细粒度任务设定,将每个RGB-D小批量视为一个任务。实验表明,该方法在有限数据下能显著改善先验(最大RMSE降低27.8%)。在元学习初始化基础上微调始终优于基线。为评估泛化性,提出零样本跨数据集协议,验证了元初始化带来的优越泛化能力,且可作为通用插件适配多种现有深度估计方法。该工作结合深度估计与元学习,推动感知技术向实际机器人应用迈进。
原文摘要 · Abstract (English)
Indoor robots rely on depth to perform tasks like navigation or obstacle detection, and single-image depth estimation is widely used to assist perception. Most indoor single-image depth prediction focuses less on model generalizability to unseen datasets, concerned with in-the-wild robustness for system deployment. This work leverages gradient-based meta-learning to gain higher generalizability on zero-shot cross-dataset inference. Unlike the most-studied meta-learning of image classification associated with explicit class labels, no explicit task boundaries exist for continuous depth values tied to highly varying indoor environments regarding object arrangement and scene composition. We propose fine-grained task that treats each RGB-D mini-batch as a task in our meta-learning formulation. We first show that our method on limited data induces a much better prior (max 27.8% in RMSE). Then, finetuning on meta-learned initialization consistently outperforms baselines without the meta approach. Aiming at generalization, we propose zero-shot cross-dataset protocols and validate higher generalizability induced by our meta-initialization, as a simple and useful plugin to many existing depth estimation methods. The work at the intersection of depth and meta-learning potentially drives both research to step closer to practical robotic and machine perception usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。