无需逐帧标注,用粗粒度标签定位机器人执行失败动作
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

- 将失败检测设为粗监督学习,利用轨迹级标签定位故障点
- 在三个数据集和真实机器人上实现最佳多任务检测效果
- 适合需要实时可靠性监控的具身智能系统部署
视觉-语言-动作(VLA)模型使机器人能理解自然语言指令并泛化于多样任务,但在实际部署中仍易发生执行失败,影响可靠性。现有方法或依赖昂贵的动作重采样,或需外部模型,而其他方法将轨迹级标签均匀传播至每一步,掩盖了局部故障信号。本文提出「Hide-and-Seek」框架,将VLA失败检测建模为粗粒度监督学习问题。通过结合跨轨迹与内轨迹对比目标,该方法仅凭轨迹级标注即可定位故障动作,并生成具有时间结构的失败信号。我们在LIBERO、VLABench及真实机器人平台上的三种代表性VLA策略(OpenVLA、π₀、π₀.₅)上进行评估,结果表明,该方法在符合预测下实现了领先的多任务失败检测性能,兼具实用性准确率与时效性权衡,并可良好泛化至已见与未见任务。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during execution is therefore critical for the robust deployment of embodied systems. Existing failure detection methods either rely on expensive action resampling or external models, while alternatives propagate trajectory-level labels uniformly across every timestep, obscuring localized failure signals. In this paper, we propose \textbf{Hide-and-Seek}, a framework that formulates VLA failure detection as a coarsely supervised learning problem. By combining inter-trajectory and intra-trajectory contrastive objectives, Hide-and-Seek localizes failure-indicative actions and induces temporally structured failure signals from trajectory-level supervision alone, without any step-level annotation. We evaluate Hide-and-Seek on LIBERO, VLABench, and a real-world robotic platform across three representative VLA policies: OpenVLA, $π_0$, and $π_{0.5}$.Our method achieves state-of-the-art multi-task failure detection performance with a practical accuracy--timeliness trade-off under conformal prediction, and generalizes well to both seen and unseen tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。