通过智能视觉观察提升低成本机器人跨环境操作鲁棒性
SEVO: Semantic-Enhanced Virtual Observation for Robust VLA Manipulation via Active Illumination and Data-Centric Collection

- 用多视角相机+红光照明+实时语义分割增强视觉输入
- 训练环境成功率95%(ACT),新环境仍达85%
- 强调数据多样性比模型大小更关键,适合家庭服务机器人
视觉-语言-动作(VLA)与模仿学习策略在低成本硬件上训练后,常因部署环境变化而失效。现有评估显示在固定背景下的成功率达高,但实际应用中转移性能接近零。我们提出SEVO(语义增强虚拟观察),一种以数据为中心的方法,在不修改策略架构的前提下提升跨环境操作鲁棒性。SEVO通过三种机制改造原始RGB图像流:(1) 固定于机械臂的多视角相机覆盖完整操作空间,(2) 主动红光照明实现物体外观物理归一化,(3) 实时YOLO分割叠加提供背景无关的语义提示。关键发现:系统性地在遥操作中变化光照、背景和干扰物的数据采集协议是泛化性的最重要因素。针对透明水瓶这一易混淆物体,设计简单抓取放置任务,在两台移动平台完成数百次受控真实机器人实验。完整流程在训练环境实现ACT 95%、SmolVLA 83%的抓取成功率,迁移至新环境仍保持85%和75%。未使用SEVO时,同一策略在训练环境仅75%/70%,新环境骤降至30-35%。结果表明,合理的观测设计与数据多样性收集,而非模型规模,是使低成本机器人在日常家庭环境中可靠运行的关键。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) and imitation-learning policies trained via community toolchains on low-cost hardware frequently fail when deployed outside the training environment. Existing evaluations, including the original ACT and SmolVLA benchmarks, demonstrate high success rates under controlled, fixed backgrounds, yet community practitioners report near-zero transfer to new environments. We present SEVO (Semantic-Enhanced Virtual Observation), a data-centric approach that improves cross-environment manipulation robustness without modifying the policy architecture. SEVO transforms the raw RGB camera stream through three mechanisms: (1) body-fixed cameras whose combined fields of view cover the full manipulation workspace, (2) active red-spectrum illumination that physically normalizes object appearance, and (3) real-time YOLO segmentation overlay that provides a background-invariant semantic cue. Critically, we show that a diversified data collection protocol (systematically varying lighting, backgrounds, and distractors during teleoperation) is the single most important factor for generalization. We target transparent water bottles, objects that visually blend with their surroundings, and select a simple pick-and-place task to enable hundreds of controlled real-robot trials across two mobile platforms. The full pipeline achieves 95% grasp success with ACT and 83% with SmolVLA in the training environment, transferring to novel environments at 85% and 75%. Without SEVO, the same policies achieve only 75%/70% in training and collapse to 30-35% in novel environments. Our results demonstrate that principled observation design and environmental diversity during data collection, not model scaling, enable low-cost robots to operate reliably in everyday household environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。