arXiv:2604.15814cs.CVcs.RO2026-04被引 1

解决机器人在开放世界中持续校准手眼关系的问题

Continual Hand-Eye Calibration for Open-world Robotic Manipulation

论文配图:Continual Hand-Eye Calibration for Open-world Robotic Manipulation
图 1 · 摘自论文原文
  • 用空间感知回放策略保留历史场景代表性样本
  • 通过结构保持的双蒸馏缓解遗忘,提升多场景适应性
  • 适合需要长期适应新环境的机器人操作任务

基于视觉定位的手眼标定是开放世界环境下机器人操作的关键能力。然而,大多数基于深度学习的标定模型在面对未见数据和场景变化时易出现灾难性遗忘,而简单的回放式持续学习策略难以有效缓解此问题。为此,本文提出一种持续手眼标定框架,通过空间感知回放策略(SARS)和结构保持型双蒸馏(SPDD)实现对连续出现的开放世界操作场景的自适应。SARS构建几何均匀的回放缓冲区,以最具信息量的视角替代冗余相邻帧,确保各场景姿态空间的全面覆盖;SPDD将定位知识分解为粗粒度场景布局与细粒度位姿精度,并分别进行蒸馏,有效缓解两类遗忘。当新场景到来时,SARS从所有先前场景中提供几何代表性回放样本,SPDD在此基础上实施结构化蒸馏以保留已有知识。新场景训练完成后,将精选样本加入回放缓冲区,支持未来持续学习。多个公开数据集上的实验表明,该框架显著提升了抗场景遗忘性能,在保持旧场景精度的同时仍能有效适应新场景,验证了其有效性。

原文摘要 · Abstract (English)

Hand-eye calibration through visual localization is a critical capability for robotic manipulation in open-world environments. However, most deep learning-based calibration models suffer from catastrophic forgetting when adapting into unseen data amongst open-world scene changes, while simple rehearsal-based continual learning strategy cannot well mitigate this issue. To overcome this challenge, we propose a continual hand-eye calibration framework, enabling robots to adapt to sequentially encountered open-world manipulation scenes through spatially replay strategy and structure-preserving distillation. Specifically, a Spatial-Aware Replay Strategy (SARS) constructs a geometrically uniform replay buffer that ensures comprehensive coverage of each scene pose space, replacing redundant adjacent frames with maximally informative viewpoints. Meanwhile, a Structure-Preserving Dual Distillation (SPDD) is proposed to decompose localization knowledge into coarse scene layout and fine pose precision, and distills them separately to alleviate both types of forgetting during continual adaptation. As a new manipulation scene arrives, SARS provides geometrically representative replay samples from all prior scenes, and SPDD applies structured distillation on these samples to retain previously learned knowledge. After training on the new scene, SARS incorporates selected samples from the new scene into the replay buffer for future rehearsal, allowing the model to continuously accumulate multi-scene calibration capability. Experiments on multiple public datasets show significant anti scene forgetting performance, maintaining accuracy on past scenes while preserving adaptation to new scenes, confirming the effectiveness of the framework.

机器人操作持续学习手眼标定视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。