通过任务感知的视觉表示,让机器人在外观变化下仍能稳定执行动作。
Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy
- 用分割和重绘生成统一语义-几何图像,消除无关视觉干扰。
- 在多个基准上保持原性能的同时,显著提升对外观变化的鲁棒性。
- 适合需要高视觉鲁棒性的机器人操控任务,如抓取与操作。
从示范中学习的视觉-运动策略常因原始RGB观测中的无关视觉因素而过拟合,导致背景变化或物体变色时行为脆弱。本文提出一种任务感知的观察接口,将视觉输入归一化为共享表征,在不修改或微调策略的情况下提升对分布外(OOD)外观变化的鲁棒性。给定RGB图像和开放词汇的任务相关实体描述,使用SAM3分割目标物体与机械臂/夹爪,构建L0观测:将分割区域用预设语义色重绘于固定背景上。对于需要更强几何信息的任务,进一步通过深度引导覆盖,将Depth Anything 3的单目深度注入分割区域,生成统一的语义-几何观测(L1),仍为标准三通道图像输入。在RoboMimic(Lift)、ManiSkill YCB杂乱场景抓取、四个RLBench可控外观变化任务及两个真实世界Franka任务(ReachX与CloseCabinet)上评估。无论使用流匹配策略还是SmolVLA模型,该接口均保持分布内性能,同时大幅增强对分布外视觉变化的鲁棒性。
原文摘要 · Abstract (English)
Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as background changes and object recoloring. We propose a task-aware observation interface that canonicalizes visual input into a shared representation, improving robustness to out-of-distribution (OOD) appearance changes without modifying or fine-tuning the policy. Given an RGB image and an open-vocabulary specification of task-relevant entities, we use SAM3 to segment the target object and robot/gripper. We construct an L0 observation by repainting segmented entities with predefined semantic colors on a constant background. For tasks requiring stronger geometric cues, we further inject monocular depth from Depth Anything 3 into the segmented regions via depth-guided overwrite, yielding a unified semantic--geometric observation (L1) that remains a standard 3-channel, image-like input. We evaluate on RoboMimic (Lift), ManiSkill YCB grasping under clutter, four RLBench tasks under controlled appearance shifts, and two real-world Franka tasks (ReachX and CloseCabinet). Across benchmarks and policy backbones (Flow Matching Policy and SmolVLA), our interface preserves in-distribution performance while substantially improving robustness under OOD visual shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。