arXiv:2503.09632cs.ROcs.CV2025-03

用扩散模型修复视觉异常,提升机器人遥操作稳定性。

Adaptive Anomaly Recovery for Telemanipulation: A Diffusion Model Approach to Vision-Based Tracking

  • 通过帧差检测识别视频异常段,用扩散模型重建
  • 在不同遮挡时长下,均方误差降低17.2%~51.1%
  • 适合高维视觉数据的连续控制任务,如远程手术机器人

灵巧的遥操作依赖于对操作者指令的持续稳定追踪。基于视觉的追踪方法虽广泛使用,但易受遮挡、光照不足或视线丢失等异常影响。传统滤波、回归与插值方法仅处理低维数据(如角度和位置),常导致信息损失。近期基于扩散模型的方法在高维视频重建与生成中表现优异,但尚未应用于机器人连续控制任务。本文提出扩散增强型遥操作框架(DET),结合帧差检测(FDD)识别视频流中的异常片段,并利用扩散模型重建,从而保障复杂视觉条件下的操作鲁棒性。在多种异常场景下验证,相比三次样条插值,平均均方误差降低17.2%;相比基于FFT的插值,降低51.1%。

原文摘要 · Abstract (English)

Dexterous telemanipulation critically relies on the continuous and stable tracking of the human operator's commands to ensure robust operation. Vison-based tracking methods are widely used but have low stability due to anomalies such as occlusions, inadequate lighting, and loss of sight. Traditional filtering, regression, and interpolation methods are commonly used to compensate for explicit information such as angles and positions. These approaches are restricted to low-dimensional data and often result in information loss compared to the original high-dimensional image and video data. Recent advances in diffusion-based approaches, which can operate on high-dimensional data, have achieved remarkable success in video reconstruction and generation. However, these methods have not been fully explored in continuous control tasks in robotics. This work introduces the Diffusion-Enhanced Telemanipulation (DET) framework, which incorporates the Frame-Difference Detection (FDD) technique to identify and segment anomalies in video streams. These anomalous clips are replaced after reconstruction using diffusion models, ensuring robust telemanipulation performance under challenging visual conditions. We validated this approach in various anomaly scenarios and compared it with the baseline methods. Experiments show that DET achieves an average RMSE reduction of 17.2% compared to the cubic spline and 51.1% compared to FFT-based interpolation for different occlusion durations.

遥操作扩散模型视觉异常机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。