提升机器人操作模型在复杂场景下的空间泛化能力
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
- 融合视觉特征与空间表征,增强政策优化中的空间感知
- 设计几何进度驱动的密集奖励,提升训练稳定性
- 适合需要强空间推理的机器人操控任务研究者
视觉-语言-动作(VLA)模型在机器人操作中表现出强大泛化能力,但强化学习微调常因空间分布变化导致鲁棒性下降。对于基于流匹配的VLA策略,这种退化与强化学习适配过程中空间归纳偏置的削弱密切相关,稀疏奖励和空间无关探索逐渐偏向短时程视觉线索。为此,我们提出SA-VLA——一种空间感知的强化学习适应框架,通过将表征学习、奖励设计与探索策略与任务几何结构对齐,保持政策优化中的空间接地性。SA-VLA融合隐式空间表征与视觉标记,提供反映几何进展的密集奖励,并采用针对流匹配动态设计的扫描式空间条件退火探索策略(SCAN)。在多物体及杂乱操作基准测试中,SA-VLA实现稳定强化学习微调,显著提升零样本空间泛化性能,生成更鲁棒且可迁移的行为。代码与项目页面见https://xupan.top/Projects/savla。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models exhibit strong generalization in robotic manipulation, yet reinforcement learning (RL) fine-tuning often degrades robustness under spatial distribution shifts. For flow-matching VLA policies, this degradation is closely associated with the erosion of spatial inductive bias during RL adaptation, as sparse rewards and spatially agnostic exploration increasingly favor short-horizon visual cues. To address this issue, we propose \textbf{SA-VLA}, a spatially-aware RL adaptation framework that preserves spatial grounding during policy optimization by aligning representation learning, reward design, and exploration with task geometry. SA-VLA fuses implicit spatial representations with visual tokens, provides dense rewards that reflect geometric progress, and employs \textbf{SCAN}, a spatially-conditioned annealed exploration strategy tailored to flow-matching dynamics. Across challenging multi-object and cluttered manipulation benchmarks, SA-VLA enables stable RL fine-tuning and improves zero-shot spatial generalization, yielding more robust and transferable behaviors. Code and project page are available at https://xupan.top/Projects/savla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。