arXiv:2604.13788cs.ROcs.CV2026-04中稿 · ICRA被引 1

用视觉语言模型区分机器人异常中的真实故障与正常偏差

Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

论文配图:Failure Identification in Imitation Learning Via Statistical and Semantic Filtering
图 1 · 摘自论文原文
  • 通过最优传输匹配构建正常演示的紧凑表征,生成异常评分和热图
  • 在真实数据集上比现有方法失败检测准确率提升17.38%
  • 适合需要高可靠性部署的工业机器人场景

模仿学习(IL)在受控环境下表现良好,但在实际部署中仍易因罕见事件(如硬件故障、缺陷零件、意外人为动作或训练分布外状态)导致执行失败。基于视觉的异常检测(AD)方法虽能识别异常状态,却无法区分故障与良性偏离。本文提出FIDeL(Failure Identification in Demonstration Learning),一种独立于策略的故障检测模块。该模块结合最新AD方法,构建正常示范的紧凑表征,并通过最优传输匹配对输入观测进行对齐,生成异常分数与热图。利用扩展的符合性预测法确定时空阈值,并引入视觉-语言模型(VLM)进行语义过滤,以区分良性异常与真实故障。此外,本文还构建了多模态真实任务数据集BotFails,用于机器人故障检测研究。FIDeL在多个基准测试中优于现有方法,在异常检测上达到+5.30% AUROC提升,在BotFails上失败检测准确率提高+17.38%。

原文摘要 · Abstract (English)

Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare events such as hardware faults, defective parts, unexpected human actions, or any state that lies outside the training distribution can lead to failed executions. Vision-based Anomaly Detection (AD) methods emerged as an appropriate solution to detect these anomalous failure states but do not distinguish failures from benign deviations. We introduce FIDeL (Failure Identification in Demonstration Learning), a policy-independent failure detection module. Leveraging recent AD methods, FIDeL builds a compact representation of nominal demonstrations and aligns incoming observations via optimal transport matching to produce anomaly scores and heatmaps. Spatio-temporal thresholds are derived with an extension of conformal prediction, and a Vision-Language Model (VLM) performs semantic filtering to discriminate benign anomalies from genuine failures. We also introduce BotFails, a multimodal dataset of real-world tasks for failure detection in robotics. FIDeL consistently outperforms state-of-the-art baselines, yielding +5.30% percent AUROC in anomaly detection and +17.38% percent failure-detection accuracy on BotFails compared to existing methods.

机器人异常检测模仿学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。