arXiv:2511.22697cs.ROcs.CL2025-11被引 13

用少量示范精准调整视觉语言动作模型,提升机器人任务适应性

Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations

  • 基于机制可解释性,根据任务特性选择性微调注意力头
  • 在真实机械臂上测试,性能优于LoRA且计算开销更低
  • 适合需要快速适配新任务的机器人研究者使用

视觉-语言-动作(VLAs)模型有望将视觉-语言模型的成功拓展至机器人领域。然而,与视觉-语言领域不同,机器人任务中的VLAs需通过微调应对机器人本体、环境特征及任务空间关系等物理因素差异。现有微调方法缺乏针对性,对所有任务采用相同的参数调整。受神经科学中功能特异性的启发,我们提出假设:针对特定任务微调稀疏模型表征更有效。本文提出Robotic Steering方法,基于机制可解释性,利用少量示范识别并选择性微调与任务的物理、视觉和语言需求匹配的注意力头。在Franka Emika机械臂上的全面实机评估表明,该方法性能超越LoRA,具备更强的任务变化鲁棒性、更低的计算成本以及更高的可解释性,能高效适配多样化机器人任务。

原文摘要 · Abstract (English)

Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for robotics require finetuning to contend with varying physical factors like robot embodiment, environment characteristics, and spatial relationships of each task. Existing fine-tuning methods lack specificity, adapting the same set of parameters regardless of a task's visual, linguistic, and physical characteristics. Inspired by functional specificity in neuroscience, we hypothesize that it is more effective to finetune sparse model representations specific to a given task. In this work, we introduce Robotic Steering, a finetuning approach grounded in mechanistic interpretability that leverages few-shot demonstrations to identify and selectively finetune task-specific attention heads aligned with the physical, visual, and linguistic requirements of robotic tasks. Through comprehensive on-robot evaluations with a Franka Emika robot arm, we demonstrate that Robotic Steering outperforms LoRA while achieving superior robustness under task variation, reduced computational cost, and enhanced interpretability for adapting VLAs to diverse robotic tasks.

机器人微调可解释性视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。