让视觉模型在动作训练中保留关键视觉结构,提升机器人任务成功率。
FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning

- 通过在线探测动作特征方向,过滤视觉表征中的冗余信息。
- 在六种仿真和真实任务中,成功率最高提升10.7个百分点。
- 适合关注机器人视觉-语言-动作对齐的开发者与研究者。
视觉-语言-动作(VLA)策略的动作监督微调虽能有效拟合示范,但仅约束预测动作的变化方向,导致动作等价状态间的视觉结构易坍缩。本文将此现象定义为局部动作纤维上的视觉残差坍缩,并提出FiberTune训练目标,在不增加推理开销的前提下,保留教师模型结构化的视觉残差。FiberTune利用在线动作探测器估计动作可预测特征方向,从中间视觉标记表示中滤除这些方向,同时对剩余残差进行对齐与有效秩正则化。在相同训练条件下,相比仅使用任务损失的微调,FiberTune在涵盖两个基准、两种架构(pi_0.5 和 OpenVLA-OFT)的六种受控仿真设置中均表现更优,物理机器人SO-101拾取放置任务的成功率从72.7%提升至78.1%,长时程CALVIN ABC-to-D任务成功率达+10.7个百分点。残差诊断显示,性能提升与探针过滤后的残差与教师模型对齐度及有效秩增强一致。
原文摘要 · Abstract (English)
Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but constrains only the directions that change predicted actions, leaving visual structure consistent across action-equivalent states free to collapse. We formalize this as residual visual collapse along local action fibers and propose FiberTune, a training-time objective that preserves teacher-structured visual residuals without adding inference-time overhead. FiberTune uses an online action probe to estimate action-predictive feature directions, filters them from intermediate visual-token representations, and aligns the resulting probe-filtered residuals to a frozen visual teacher while regularizing their effective rank. Under identical training conditions, FiberTune improves over task-loss-only fine-tuning in every one of six controlled simulation settings spanning two benchmarks and two architectures (pi_0.5 and OpenVLA-OFT), as well as on physical SO-101 pick-place; representative gains include +10.7 percentage points SR(5) on long-horizon CALVIN ABC-to-D and physical SO-101 task success rising from 72.7% to 78.1%. Residual diagnostics show that these gains coincide with increased probe-filtered residual teacher alignment and effective rank, consistent with the action-fiber motivation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。