用3D高斯点云实现视角自适应,让冻结的视觉语言动作模型更抗视角变化。
GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting

- 用3D高斯点云生成新视角图像,自动校正部署时的相机偏移。
- 在LIBERO基准上,视角偏移后成功率从约10%提升至接近90%。
- 无需重训练模型,适配多种架构、任务和扰动程度,轻量高效。
本文提出一种轻量级、即插即用的框架,提升视觉-语言-动作(VLA)策略对视角变化的鲁棒性,无需重新训练策略。据我们所知,这是首个直接利用基于3D高斯的新视角合成技术,用于VLA策略的观测空间适配的方法。当前VLA性能依赖于训练与部署相机配置一致的隐含假设。实验表明,仅摄像头支架微小位移即可使LIBERO基准上的成功率从约90%降至约10%。此前方法如大规模微调或生成数据增强,计算成本高且易导致灾难性遗忘。为此,视角偏移被重新建模为局部新视角合成问题。在局部性假设下——相机扰动保持在工作区附近小范围内——视角归一化退化为与场景和策略无关的遮挡填补任务。本工作通过一个400万参数的3D高斯归一化模块,前置于冻结的VLA策略前实现该思想。无需修改策略权重,GS-VLA在三个独立维度上均提升性能:(1) 不同策略架构,(2) 未见过的任务套件,(3) 不同扰动规模。结果表明,轻量视觉模块可恢复视角偏移中丢失的大部分性能,且无需重训练。
原文摘要 · Abstract (English)
This paper proposes a lightweight, plug-and-play framework that improves robustness to viewpoint shifts in Vision-Language-Action (VLA) policies without policy retraining. To our knowledge, this is the first approach to directly leverage 3D Gaussian-based novel-view synthesis for observation-space adaptation in VLA policies. Current VLA performance relies on the implicit assumption that training and deployment camera configurations are identical. Our experiments show that even a small displacement of the camera mount can reduce the success rate on the LIBERO benchmark from about 90% to about 10% in the worst case. Prior approaches, such as large-scale fine-tuning or generative data augmentation, are computationally expensive and risk catastrophic forgetting. To address this, viewpoint shifts are reformulated as a localized novel-view synthesis problem. Under a Locality assumption, that camera perturbations remain within a small bounded region relative to the workspace, viewpoint normalization reduces to a scene- and policy-independent disocclusion task. Our work implements this idea with a 4M-parameter 3D-Gaussian canonicalizer prepended to a frozen VLA policy. Without modifying policy weights, GS-VLA improves performance across three orthogonal axes: (1) Policy architectures, (2) Unseen task suites, and (3) Perturbation scales. These results show that a lightweight visual module can recover a large fraction of the performance lost under viewpoint shift, without policy retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。