arXiv:2609.04277cs.ROcs.CV2026-09

用弱监督和主动学习实现低成本的机器人动作失败精准定位。

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

论文配图:FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 从无标签动作片段中自动提取异常模式作为弱监督信号。
  • 仅对最不确定轨迹进行标注,显著降低人工成本。
  • 在多个机器人模型上实现更准的失败检测与时间点定位。

视觉-语言-动作(VLA)策略在通用机器人操作中展现出巨大潜力,但在长周期执行中仍可能意外失败,因此可靠失败检测对安全部署至关重要。现有方法要么依赖视觉模型,在错误动作发生后才检测失败,要么使用基于VLA内部表示的轻量级主动检测器,但这些方法通常依赖轨迹级标签,导致失败轨迹中正常的前期行为被误标为失败,造成标签噪声,限制了轨迹级检测准确性和精确的时间戳定位。本文研究细粒度的时间戳级失败检测,同时应对密集标注的成本问题。提出一种数据高效的框架:首先利用无标签的VLA动作片段构建动作衍生的弱监督信号,捕捉不一致连续片段、停滞或空闲动作、剧烈随机运动等异常模式;然后通过主动学习选择最不确定的轨迹进行时间戳级标注,并用这些信息丰富的标签微调检测器。在多个VLA策略上的实验表明,该方法提升了时间戳级和轨迹级失败检测性能。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies have shown strong potential for general-purpose robotic manipulation, but they can still fail unpredictably during long-horizon execution, making reliable failure detection essential for safe deployment. Existing methods either rely on visual models that typically detect failures only after erroneous actions have occurred, or use lightweight proactive detectors trained on VLA internal representations. However, these proactive methods are often supervised with trajectory-level labels, causing normal pre-failure behavior in unsuccessful trajectories to be incorrectly labeled as failure. This supervision mismatch introduces label noise and limits both trajectory-level detection accuracy and precise timestamp-level failure localization. In this work, we study fine-grained timestamp-level VLA failure detection while addressing the cost of dense annotation. We propose a data-efficient framework that first leverages unlabeled VLA action chunks to construct action-derived weak supervision signals, capturing abnormal patterns such as inconsistent consecutive chunks, frozen or idle actions, and aggressive random motions. We then use active learning to select only the most uncertain trajectories for timestamp-level annotation and fine-tune the detector with these informative labels. Experiments across multiple VLA policies show that our method improves both timestamp-level and trajectory-level failure detection performance.

机器人失败检测主动学习弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。