arXiv:2601.18923cs.ROcs.CV2026-01被引 3

为机器人深度图像训练首个自监督基础模型,实现跨场景泛化。

DeFM: Learning Foundation Representations from Depth for Robotics

  • 基于6000万张深度图,用DINO式自蒸馏学习几何语义特征
  • 在分类、分割、导航等任务上达最先进水平,支持真实世界部署
  • 提供轻量化模型,适合资源受限的机器人系统直接使用

深度传感器广泛部署于机器人平台,快速高保真深度模拟的进步使基于深度观测的机器人策略在多种任务中实现了可靠的仿真到现实迁移。尽管如此,与已由大规模基础模型主导的RGB模态相比,深度模态的表征学习仍处于探索阶段。为此,我们提出DeFM,一个完全基于深度图像训练的自监督基础模型,专用于机器人应用。在包含6000万张深度图像的精选数据集上,采用DINO式自蒸馏目标,DeFM学习到可泛化至多样环境、任务和传感器的几何与语义表示。为保持多尺度度量感知,我们引入一种新颖的输入归一化策略。进一步将DeFM压缩为适合资源受限机器人的紧凑模型。在基于深度的分类、分割、导航、运动和操作基准测试中,DeFM达到最先进性能,并展现出从仿真到真实环境的强大泛化能力。我们发布所有预训练模型,可无需任务微调直接用于深度驱动的机器人学习。

原文摘要 · Abstract (English)

Depth sensors are widely deployed across robotic platforms, and advances in fast, high-fidelity depth simulation have enabled robotic policies trained on depth observations to achieve robust sim-to-real transfer for a wide range of tasks. Despite this, representation learning for depth modality remains underexplored compared to RGB, where large-scale foundation models now define the state of the art. To address this gap, we present DeFM, a self-supervised foundation model trained entirely on depth images for robotic applications. Using a DINO-style self-distillation objective on a curated dataset of 60M depth images, DeFM learns geometric and semantic representations that generalize to diverse environments, tasks, and sensors. To retain metric awareness across multiple scales, we introduce a novel input normalization strategy. We further distill DeFM into compact models suitable for resource-constrained robotic systems. When evaluated on depth-based classification, segmentation, navigation, locomotion, and manipulation benchmarks, DeFM achieves state-of-the-art performance and demonstrates strong generalization from simulation to real-world environments. We release all our pretrained models, which can be adopted off-the-shelf for depth-based robotic learning without task-specific fine-tuning. Webpage: https://de-fm.github.io/

深度学习机器人自监督基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。