用视觉物理双校验提升机器人数据质量,让大模型更靠谱。
VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

- 构建鱼眼视角VQA数据集,对齐扭曲图像与语言理解
- 通过物理合规性评分过滤无效轨迹,提升训练数据可信度
- 联合训练视觉语言与动作预测,显著提升真实场景成功率
通用操控接口(UMI)实现了无需特定硬件的规模化机器人数据采集,但将此类数据用于训练大规模视觉-语言-动作(VLA)模型仍面临根本挑战。我们识别出两大核心不匹配:腕装鱼眼相机视角存在严重径向畸变且以夹爪为中心,超出预训练视觉语言模型(VLM)分布;人类采集的轨迹常违反运动学极限、发生碰撞或超出控制器带宽,导致训练出物理不可行的动作。为此,我们提出VISTA框架,通过三个协同组件弥合双重差距:(i) UMI-VQA,首个面向腕装鱼眼观测的大规模VQA数据集,通过辅助视觉-语言监督使VLM表征适应畸变视觉环境;(ii) 系统性物理验证流程,在数据进入训练前进行完整性检查,并评估每条轨迹的连续性、自碰撞风险与执行保真度;(iii) 两阶段联合训练策略,同时在UMI-VQA上学习视觉语言对齐,在经验证的轨迹上学习动作预测。实验表明,引入UMI-VQA能持续提升下游策略性能,且物理验证得分与部署成功率强相关。在多种仿真与真实操作任务中,VISTA显著优于π_{0.5}、LingBot-VLA和Wall-X等强基线。我们开源了物理验证流程、UMI-VQA数据集、经验证的轨迹数据及预训练模型。
原文摘要 · Abstract (English)
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions. To address the challenges, we present VISTA, a framework that bridges this dual gap through three synergistic components. (i)~UMI-VQA, the first large-scale VQA dataset tailored to wrist-mounted fisheye observations, aligns VLM representations to the distorted visual regime via auxiliary vision-language supervision. (ii)~A systematic physical-validation pipeline performs a data-completeness pre-check and scores each valid trajectory for trajectory continuity, self-collision risk, and execution fidelity before it enters training. (iii)~A two-stage co-training recipe jointly learns vision-language grounding on UMI-VQA and action prediction on validated trajectories. Our experiments empirically show that incorporating UMI-VQA consistently improves downstream policy performance, and that physical-validation scores are strongly predictive of deployment success. On diverse simulation and real-world manipulation tasks, VISTA significantly outperforms strong baselines including $π_{0.5}$, LingBot-VLA, and Wall-X. We release the physical-validation pipeline, UMI-VQA, validated trajectory data, and the pre-trained model for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。