arXiv:2601.00359cs.CVcs.RO2026-01被引 1

用知识蒸馏让机器人实时理解环境,支持自然语言查询。

Efficient Prediction of Dense Visual Embeddings via Distillation and RGB-D Transformers

  • 用大模型指导小模型学习像素级视觉嵌入
  • 全模型达26.3帧/秒,小模型77.0帧/秒
  • 适合移动机器人3D建图与语言交互

在家庭环境中,机器人需全面理解周围场景才能与未训练的人类有效、直观地互动。本文提出DVEFormer——一种基于RGB-D Transformer的高效方法,通过知识蒸馏预测密集文本对齐视觉嵌入(DVE)。不同于固定类别进行传统语义分割,本方法利用Alpha-CLIP的教师嵌入引导高效的学生模型DVEFormer学习细粒度像素级嵌入。该方法仍支持经典语义分割(如线性探测),更可实现灵活的文本查询及其他应用,如构建完整3D地图。在常见室内数据集上的评估表明,本方法性能相当且满足实时要求:全模型运行在26.3 FPS,小模型达77.0 FPS(NVIDIA Jetson AGX Orin)。定性结果展示了其在真实场景中的有效性与潜在应用。整体而言,本方法可作为传统分割方案的即插即用替代,支持自然语言查询并无缝集成至移动机器人3D映射流程。

原文摘要 · Abstract (English)

In domestic environments, robots require a comprehensive understanding of their surroundings to interact effectively and intuitively with untrained humans. In this paper, we propose DVEFormer - an efficient RGB-D Transformer-based approach that predicts dense text-aligned visual embeddings (DVE) via knowledge distillation. Instead of directly performing classical semantic segmentation with fixed predefined classes, our method uses teacher embeddings from Alpha-CLIP to guide our efficient student model DVEFormer in learning fine-grained pixel-wise embeddings. While this approach still enables classical semantic segmentation, e.g., via linear probing, it further enables flexible text-based querying and other applications, such as creating comprehensive 3D maps. Evaluations on common indoor datasets demonstrate that our approach achieves competitive performance while meeting real-time requirements, operating at 26.3 FPS for the full model and 77.0 FPS for a smaller variant on an NVIDIA Jetson AGX Orin. Additionally, we show qualitative results that highlight the effectiveness and possible use cases in real-world applications. Overall, our method serves as a drop-in replacement for traditional segmentation approaches while enabling flexible natural-language querying and seamless integration into 3D mapping pipelines for mobile robotics.

视觉嵌入机器人感知知识蒸馏3D建图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。