arXiv:2604.09326cs.ROcs.CV2026-04

用多模态特征重建提升人机协作中的异常检测能力

Multimodal Anomaly Detection for Human-Robot Interaction

  • 将视频转为语义特征向量,再做重建检测
  • 融合视觉、机器人传感器与场景图,准确率显著提升
  • 适合关注人机安全的机器人研发人员

确保人机协作(HRI)中的安全与可靠性,需及时检测可能导致系统故障或危险行为的异常事件。异常检测在此过程中至关重要,使机器人能在协作任务中识别并响应运行偏离正常状态的情况。尽管重建模型已在HRI中被广泛研究,但直接作用于特征向量的方法仍较少被探索。本文提出MADRI框架,先将视频流转化为语义明确的特征向量,再进行基于重建的异常检测。同时,我们通过引入机器人的内部传感器读数和场景图(Scene Graph),使模型能够捕捉外部视觉环境异常及机器人自身内部故障。为评估该方法,我们收集了一个自定义数据集,包含在正常与异常条件下执行的简单抓取-放置任务。实验表明,仅基于视觉特征向量的重建已能有效检测异常,而融合多模态信息进一步提升了检测性能,验证了多模态特征重建在人机协作中实现鲁棒异常检测的优势。

原文摘要 · Abstract (English)

Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures or unsafe behaviours. Anomaly detection thus plays a critical role in enabling robots to recognize and respond to deviations from normal operation during collaborative tasks. While reconstruction models have been actively explored in HRI, approaches that operate directly on feature vectors remain largely unexplored. In this work, we propose MADRI, a framework that first transforms video streams into semantically meaningful feature vectors before performing reconstruction-based anomaly detection. Additionally, we augment these visual feature vectors with the robot's internal sensors' readings and a Scene Graph, enabling the model to capture both external anomalies in the visual environment and internal failures within the robot itself. To evaluate our approach, we collected a custom dataset consisting of a simple pick-and-place robotic task under normal and anomalous conditions. Experimental results demonstrate that reconstruction on vision-based feature vectors alone is effective for detecting anomalies, while incorporating other modalities further improves detection performance, highlighting the benefits of multimodal feature reconstruction for robust anomaly detection in human-robot collaboration.

异常检测人机交互多模态机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。