arXiv:2604.23001cs.ROcs.AI2026-04中稿 · TMLR after peer-re…综述被引 4

VLA机器人研究的瓶颈不在模型,而在数据基础设施。

Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines

论文配图:Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
图 1 · 摘自论文原文
  • 从数据集、评测基准到生成引擎,系统分析VLA数据链路
  • 发现真实与合成数据存在保真度-成本权衡难题
  • 适合关注机器人具身学习数据构建的研究者

尽管视觉-语言-动作(VLA)模型取得显著进展,但其核心瓶颈——支撑具身学习的数据基础设施仍被忽视。本文主张,未来VLA的突破将更多依赖于高质量数据引擎与结构化评估协议的协同设计。为此,我们围绕数据集、评测基准和数据引擎三个支柱展开系统性、以数据为中心的分析。在数据集方面,按具身多样性、模态构成和动作空间形式对真实世界与合成数据进行分类,揭示出大规模采集中普遍存在的保真度-成本权衡。在评测基准方面,联合分析任务复杂度与环境结构,暴露现有协议在组合泛化与长时程推理评估方面的结构性缺失。在数据引擎方面,考察基于仿真、视频重建与自动化任务生成三种范式,发现其均存在物理真实性不足及仿真到现实迁移困难的问题。综合分析后,提炼出四大开放挑战:表征对齐、多模态监督、推理评估与可扩展数据生成。我们强调,必须将数据基础设施视为首要科研问题,而非背景性考量。

原文摘要 · Abstract (English)

Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this survey, we argue that future advances in VLA will depend less on model architecture and more on the co-design of high-fidelity data engines and structured evaluation protocols. To this end, we present a systematic, data-centric analysis of VLA research organized around three pillars: datasets, benchmarks, and data engines. For datasets, we categorize real-world and synthetic corpora along embodiment diversity, modality composition, and action space formulation, revealing a persistent fidelity-cost trade-off that fundamentally constrains large-scale collection. For benchmarks, we analyze task complexity and environment structure jointly, exposing structural gaps in compositional generalization and long-horizon reasoning evaluation that existing protocols fail to address. For data engines, we examine simulation-based, video-reconstruction, and automated task-generation paradigms, identifying their shared limitations in physical grounding and sim-to-real transfer. Synthesizing these analyses, we distill four open challenges: representation alignment, multimodal supervision, reasoning assessment, and scalable data generation. Addressing them, we argue, requires treating data infrastructure as a first-class research problem rather than a background concern.

机器人学习数据引擎具身智能评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。