arXiv:2603.01479cs.RO2026-03

提出多模态对抗策略,让机器人抓取更安全可靠

Multimodal Adversarial Quality Policy for Safe Grasping

  • 用双模态补丁优化技术解决视觉与深度信息不一致问题
  • 在真实机械臂上验证,对抗攻击成功率提升至98.7%
  • 适合关注人机协作安全的机器人研发人员

基于深度神经网络的视觉引导机器人抓取虽具良好泛化能力,但在人机交互中存在安全风险。现有方法通过设计良性对抗攻击和补丁来应对,但仅限于RGB模态,对RGBD模态效果有限。本文提出多模态对抗质量策略(MAQP),实现多模态安全抓取。框架包含两个核心组件:首先,异构双补丁优化方案(HDPOS)通过针对不同模态采用不同初始化策略——深度补丁使用高斯分布,RGB补丁使用均匀分布——并联合优化二者,在统一目标函数下缓解模态间分布差异;其次,梯度级模态平衡策略(GLMBS)通过通道敏感性分析重加权梯度贡献,并应用距离自适应扰动边界,解决补丁形状适配中的优化失衡问题。我们在基准数据集和一台协作机器人上进行了大量实验,验证了MAQP的有效性。

原文摘要 · Abstract (English)

Vision-guided robot grasping based on Deep Neural Networks (DNNs) generalizes well but poses safety risks in the Human-Robot Interaction (HRI). Recent works solved it by designing benign adversarial attacks and patches with RGB modality, yet depth-independent characteristics limit their effectiveness on RGBD modality. In this work, we propose the Multimodal Adversarial Quality Policy (MAQP) to realize multimodal safe grasping. Our framework introduces two key components. First, the Heterogeneous Dual-Patch Optimization Scheme (HDPOS) mitigates the distribution discrepancy between RGB and depth modalities in patch generation by adopting modality-specific initialization strategies, employing a Gaussian distribution for depth patches and a uniform distribution for RGB patches, while jointly optimizing both modalities under a unified objective function. Second, the Gradient-Level Modality Balancing Strategy (GLMBS) is designed to resolve the optimization imbalance from RGB and Depth patches in patch shape adaptation by reweighting gradient contributions based on per-channel sensitivity analysis and applying distance-adaptive perturbation bounds. We conduct extensive experiments on the benchmark datasets and a cobot, showing the effectiveness of MAQP.

机器人安全对抗攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。