让视觉语言动作模型在运行时自动忽略干扰视觉,提升实际任务鲁棒性
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
- 运行时动态识别模型敏感区域,仅微调无关视觉部分
- 在干扰物和背景变化下,任务成功率下降减少至40%以内
- 无需修改模型即可兼容所有现成VLA,适合机器人部署场景
基于大规模互联网数据和机器人示范训练的视觉-语言-动作(VLA)模型有望成为通用机器人策略。然而,尽管训练规模大,这些模型对任务无关的视觉细节(如干扰物或背景颜色)仍十分脆弱。本文提出「自带VLA」(BYOVLA):一种运行时干预机制,能(1)动态识别模型敏感的图像区域,(2)利用自动化图像编辑工具最小化地修改无关区域,从而降低模型敏感性。该方法无需模型微调或权重访问,可适配任意现成VLA。硬件实验表明,面对干扰物和背景变化时,使用BYOVLA的VLA模型几乎保持原始性能,而未加干预时任务成功率最高下降40%。更多信息、视频与代码见:https://aasherh.github.io/byovla/
原文摘要 · Abstract (English)
Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their large-scale training, VLAs are often brittle to task-irrelevant visual details such as distractor objects or background colors. We introduce Bring Your Own VLA (BYOVLA): a run-time intervention scheme that (1) dynamically identifies regions of the input image that the model is sensitive to, and (2) minimally alters task-irrelevant regions to reduce the model's sensitivity using automated image editing tools. Our approach is compatible with any off the shelf VLA without model fine-tuning or access to the model's weights. Hardware experiments on language-instructed manipulation tasks demonstrate that BYOVLA enables state-of-the-art VLA models to nearly retain their nominal performance in the presence of distractor objects and backgrounds, which otherwise degrade task success rates by up to 40%. Website with additional information, videos, and code: https://aasherh.github.io/byovla/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。