arXiv:2604.03956cs.CVcs.AI2026-04中稿 · ACL被引 7

提出VLA-Forget框架,精准删除机器人模型的不安全行为而不伤及感知与决策能力。

VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models

论文配图:VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models
图 1 · 摘自论文原文
  • 分层联合优化:对视觉、跨模态和动作生成模块分阶段选择性编辑。
  • 遗忘效率提升10%,感知特异性保留22%,任务成功率保持9%以上。
  • 适合需安全可控的机器人应用,如家庭服务或工业协作场景。

视觉-语言-动作(VLA)模型正成为具身机器人的基础模型,但其部署带来新挑战:在不损害感知、语言理解与动作控制的前提下,移除不安全、虚假或涉隐私的行为。在OpenVLA类策略中,行为由融合的视觉编码器、跨模态投影器与语言主干共同生成,不良知识分布于感知、对齐与推理/动作各层,单一模块删减常失效;而传统针对独立视觉或语言模型的遗忘方法,在具身场景中易残留遗忘或造成性能无谓损失。本文提出VLA-Forget,一种混合遗忘框架,结合比例感知的选择性编辑(用于感知与跨模态)及层选择性推理/动作遗忘(用于保用性遗忘)。该框架通过分阶段更新视觉编码器、投影器与上层动作生成Transformer块,联合优化目标遗忘、感知保留与推理维持。在遗忘集行为探测与保留任务评估中,相比强基线,其遗忘效率提升10%,感知特异性保留22%,推理与任务成功率保持9%,量化后恢复率降低55%。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models are emerging as embodied foundation models for robotic manipulation, but their deployment introduces a new unlearning challenge: removing unsafe, spurious, or privacy-sensitive behaviors without degrading perception, language grounding, and action control. In OpenVLA-style policies, behavior is produced through a fused visual encoder, a cross-modal projector, and a language backbone that predicts tokenized robot actions, so undesirable knowledge can be distributed across perception, alignment, and reasoning/action layers rather than confined to a single module. Consequently, partial unlearning applied only to the vision stack or only to the language backbone is often insufficient, while conventional unlearning baselines designed for standalone vision or language models may leave residual forgetting or incur unnecessary utility loss in embodied settings. We propose VLA-Forget, a hybrid unlearning framework that combines ratio-aware selective editing for perception and cross-modal specificity with layer-selective reasoning/action unlearning for utility-preserving forgetting. VLA-Forget jointly optimizes three objectives: targeted forgetting, perceptual preservation, and reasoning retention, through staged updates over the visual encoder, projector, and upper action-generating transformer blocks. Across forget-set behavior probes and retain-task evaluations, VLA-Forget improves forgetting efficacy by 10%, preserves perceptual specificity by 22%, retains reasoning and task success by 9%, and reduces post-quantization recovery by 55% relative to strong unlearning baselines.

具身智能模型遗忘机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。