揭示视觉语言动作模型在开放世界中的真实表现,指出现有评估方式高估性能。
How VLAs (Really) Work In Open-World Environments

- 通过可复现性与一致性分析模型鲁棒性,评估操作安全与任务意识。
- 发现当前评估仅看最终状态,忽略过程风险,导致性能被夸大。
- 提出新评估协议,关注安全违规,适合研究真实场景部署的学者。
视觉语言动作模型(VLAs)在机器人领域广泛应用,在各类操作任务中取得显著进展。近期,它们被用于长时程任务,并在BEHAVIOR1K(B1K)等基准上进行评估,以解决复杂的家务问题。当前主流评估指标为成功率或基于不依赖进度的判据的局部得分,仅关注物体的最终状态,而忽视达成该状态的过程。本文认为,此类评估方式难以反映操作的安全性,可能夸大模型表现,削弱未来真实部署的核心挑战。为此,我们对B1K挑战中的前沿模型进行了深入分析,从可复现性、性能一致性、操作安全性、任务意识及任务失败关键因素等方面评估策略。进而提出新评估协议,以捕捉安全违规行为,更准确衡量模型在复杂交互场景中的真实表现。最后讨论现有VLAs的局限性,推动未来研究方向。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) have been extensively used in robotics applications, achieving great success in various manipulation problems. More recently, VLAs have been used in long-horizon tasks and evaluated on benchmarks, such as BEHAVIOR1K (B1K), for solving complex household chores. The common metric for measuring progress in such benchmarks is success rate or partial score based on satisfaction of progress-agnostic criteria, meaning only the final states of the objects are considered, regardless of the events that lead to such states. In this paper, we argue that using such evaluation protocols say little about safety aspects of operation and can potentially exaggerate reported performance, undermining core challenges for future real-world deployment. To this end, we conduct a thorough analysis of state-of-the-art models on the B1K Challenge and evaluate policies in terms of robustness via reproducibility and consistency of performance, safety aspects of policies operations, task awareness, and key elements leading to the incompletion of tasks. We then propose evaluation protocols to capture safety violations to better measure the true performance of the policies in more complex and interactive scenarios. At the end, we discuss the limitations of the existing VLAs and motivate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。