arXiv:2506.13867cs.RO2025-06被引 5

自动选关键点提升机器人策略鲁棒性,适应复杂视觉环境。

ATK: Automatic Task-driven Keypoint Selection for Robust Policy Learning

  • 根据任务自动挑选最相关的2D关键点作为状态表示
  • 仅用少量关键点即可保持策略性能并增强抗干扰能力
  • 适合需要跨场景迁移的机器人控制与真实世界模仿学习

视觉运动策略常因训练与评估环境间的视觉差异导致性能下降。依赖6D位姿等状态估计的策略需特定追踪且难以扩展,而原始传感器策略对微小视觉扰动敏感。本文利用图像帧中空间一致的2D关键点作为灵活的状态表示,应用于模拟到现实的迁移及真实世界模仿学习。针对不同物体和任务的关键点选择差异,提出ATK方法,实现任务驱动的关键点自动筛选,使所选关键点能有效预测最优行为。该方法优化出最小化关键点集合,聚焦任务相关区域,同时保证策略性能与鲁棒性。通过蒸馏专家数据(仿真中的专家策略或人类专家),构建基于RGB图像并跟踪选定关键点的策略。借助预训练视觉模块,系统能有效编码状态,并在存在广泛场景变化及透明物体、精细任务、可变形物体操作等感知挑战下成功将策略迁移到真实场景。我们在多种机器人任务上验证了ATK,结果表明,这种最小关键点表示显著提升了对视觉扰动和环境变化的鲁棒性。

原文摘要 · Abstract (English)

Visuomotor policies often suffer from perceptual challenges, where visual differences between training and evaluation environments degrade policy performance. Policies relying on state estimations, like 6D pose, require task-specific tracking and are difficult to scale, while raw sensor-based policies may lack robustness to small visual disturbances. In this work, we leverage 2D keypoints--spatially consistent features in the image frame--as a flexible state representation for robust policy learning and apply it to both sim-to-real transfer and real-world imitation learning. However, the choice of which keypoints to use can vary across objects and tasks. We propose a novel method, ATK, to automatically select keypoints in a task-driven manner so that the chosen keypoints are predictive of optimal behavior for the given task. Our proposal optimizes for a minimal set of keypoints that focus on task-relevant parts while preserving policy performance and robustness. We distill expert data (either from an expert policy in simulation or a human expert) into a policy that operates on RGB images while tracking the selected keypoints. By leveraging pre-trained visual modules, our system effectively encodes states and transfers policies to the real-world evaluation scenario despite wide scene variations and perceptual challenges such as transparent objects, fine-grained tasks, and deformable objects manipulation. We validate ATK on various robotic tasks, demonstrating that these minimal keypoint representations significantly improve robustness to visual disturbances and environmental variations. See all experiments and more details at https://yunchuzhang.github.io/ATK/.

机器人控制关键点选择策略鲁棒性视觉迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。