arXiv:2602.10980cs.RO2026-02被引 6

提出真实世界评估基准RADAR,检验视觉-语言-动作模型在复杂物理环境下的泛化能力。

RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation

  • 构建包含真实动态因素的评测体系,涵盖物体配置、光照变化等
  • 模型在传感器噪声下3D IoU从0.261降至0.068,性能严重退化
  • 全自动化3D评估,无需人工干预,适合大规模模型对比

VLA模型在具身智能领域取得显著进展,但其评估仍局限于仿真或高度受限的真实场景,导致显著的现实差距。现有评测存在三大系统性缺陷:(1)未建模真实世界动态,忽略物体配置、机器人初始状态、光照变化与传感器噪声等关键因素;(2)忽视空间-物理智能,评价任务仅限机械操作,无法检验几何推理能力;(3)缺乏可扩展的全自动评估机制,依赖简单2D指标或人力介入,成本高且不可靠。为此,我们提出RADAR(Real-world Autonomous Dynamics And Reasoning)基准,系统评估VLA模型在真实条件下的泛化能力。该基准包含三大核心组件:(1)规范化的物理动态模拟;(2)专门设计的空间推理与物理理解任务;(3)基于3D指标的全自动评估流程,无需人工监督。我们将RADAR应用于多个先进VLA模型,发现其表面能力下隐藏严重脆弱性。在适度物理动态下,3D IoU从0.261骤降至0.068;模型空间推理能力亦极为有限。RADAR为实现可靠、通用的真实世界评估提供了必要基准。

原文摘要 · Abstract (English)

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings. This mismatch creates a substantial reality gap, where strong benchmark performance often masks poor generalization in diverse physical environments. We identify three systemic shortcomings in current benchmarking practices that hinder fair and reliable model comparison. (1) Existing benchmarks fail to model real-world dynamics, overlooking critical factors such as dynamic object configurations, robot initial states, lighting changes, and sensor noise. (2) Current protocols neglect spatial--physical intelligence, reducing evaluation to rote manipulation tasks that do not probe geometric reasoning. (3) The field lacks scalable fully autonomous evaluation, instead relying on simplistic 2D metrics that miss 3D spatial structure or on human-in-the-loop systems that are costly, biased, and unscalable. To address these limitations, we introduce RADAR (Real-world Autonomous Dynamics And Reasoning), a benchmark designed to systematically evaluate VLA generalization under realistic conditions. RADAR integrates three core components: (1) a principled suite of physical dynamics; (2) dedicated tasks that explicitly test spatial reasoning and physical understanding; and (3) a fully autonomous evaluation pipeline based on 3D metrics, eliminating the need for human supervision. We apply RADAR to audit multiple state-of-the-art VLA models and uncover severe fragility beneath their apparent competence. Performance drops precipitously under modest physical dynamics, with the expectation of 3D IoU declining from 0.261 to 0.068 under sensor noise. Moreover, models exhibit limited spatial reasoning capability. These findings position RADAR as a necessary bench toward reliable and generalizable real-world evaluation of VLA models.

VLA模型评估基准空间推理真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。