让机器人在光照变化下仍能准确操作并给出原因解释。
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
- 用视觉-语言-动作模型识别光照干扰并生成自然语言解释。
- 对抗攻击下动作预测误差降低21.7%(当前动作)和18.4%(下一步动作)。
- 适合需要安全可靠、可解释的农业机器人系统研发者使用。
智能农业正通过自动化与智能控制推动现代农业发展。然而,依赖RGB相机感知与机械臂控制的系统易受色彩、光照、噪声等光度扰动影响,在对抗攻击下可能失效。为此,本文提出基于OpenVLA-OFT框架的可解释鲁棒视觉-语言-动作模型,集成Evidence-3模块,可检测光度扰动并生成其成因与影响的自然语言解释。实验表明,相比基线模型,该模型在对抗条件下将当前动作的L1损失降低21.7%,下一步动作的L1损失降低18.4%,显著提升动作预测精度与可解释性。
原文摘要 · Abstract (English)
Smart farming has emerged as a key technology for advancing modern agriculture through automation and intelligent control. However, systems relying on RGB cameras for perception and robotic manipulators for control, common in smart farming, are vulnerable to photometric perturbations such as hue, illumination, and noise changes, which can cause malfunction under adversarial attacks. To address this issue, we propose an explainable adversarial-robust Vision-Language-Action model based on the OpenVLA-OFT framework. The model integrates an Evidence-3 module that detects photometric perturbations and generates natural language explanations of their causes and effects. Experiments show that the proposed model reduces Current Action L1 loss by 21.7% and Next Actions L1 loss by 18.4% compared to the baseline, demonstrating improved action prediction accuracy and explainability under adversarial conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。