arXiv:2603.20679cs.RO2026-03

用全景深度信息指导单目模型,提升机器人视觉策略的泛化与效率。

Enhancing Vision-Based Policies with Omni-View and Cross-Modality Knowledge Distillation for Mobile Robots

  • 通过全景深度教师模型指导单目学生模型,融合多视角与深度信息。
  • 在真实场景中,单目策略性能超越仅模仿动作的基线方法。
  • 适合资源受限的轻量级移动机器人,尤其关注导航与任务泛化能力。

基于视觉的策略广泛应用于机器人抓取与行走等任务。然而,在轻量级移动机器人上,这类策略面临三大挑战:场景迁移能力有限、本地计算资源受限以及传感器硬件成本高。为此,本文提出一种知识蒸馏方法,将信息丰富且对外观不变的全景深度策略的知识,迁移到轻量级单目策略中。核心思想是让学生不仅模仿专家的动作,还对齐全景深度教师的潜在表征。实验表明,使用全景视图和深度输入可显著提升场景迁移与导航性能;所提出的蒸馏方法相比仅模仿动作的策略,进一步提升了单目策略的表现。真实世界实验验证了该方法的有效性与实用性。代码将公开发布。

原文摘要 · Abstract (English)

Vision-based policies are widely applied in robotics for tasks such as manipulation and locomotion. On lightweight mobile robots, however, they face a trilemma of limited scene transferability, restricted onboard computation resources, and sensor hardware cost. To address these issues, we propose a knowledge distillation approach that transfers knowledge from an information-rich, appearance invariant omniview depth policy to a lightweight monocular policy. The key idea is to train the student not only to mimic the expert actions but also to align with the latent embeddings of the omni view depth teacher. Experiments demonstrate that omni-view and depth inputs improve the scene transfer and navigation performance, and that the proposed distillation method enhances the performance of a singleview monocular policy, compared with policies solely imitating actions. Real world experiments further validate the effectiveness and practicality of our approach. Code will be released publicly.

视觉策略知识蒸馏机器人单目感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。