让机器人动作表征保持语义结构,提升泛化能力
Semantic Anchoring for Robotic Action Representations

- 用语义流形锚定动作表征,分离共享与私有信息
- 实测在真实场景下任务成功率提升18.7%,跨分布泛化提升21.5%
- 无需修改部署模型,可直接接入多种视觉-语言-动作框架
视觉-语言-动作(VLA)模型继承自预训练视觉-语言模型的丰富语义表征,但在有限机器人示范数据上微调后,这种结构会退化,损害泛化性能。核心问题在于:什么样的动作表征是好的?受镜像神经元理论启发——观察与执行共享意图层级编码,我们探究机器人动作表征是否保留了预训练编码器捕获的语义结构。系统性探测证实,该结构在微调过程中逐渐退化,且其质量与任务成功率及分布外泛化能力同步提升。我们进一步提出一种即插即用方法,将动作表征锚定于语义流形,分解为共享语义通道与私有通道,推理时全部丢弃,不改变部署模型。在多种VLA骨干网络、仿真与真实世界基准上验证,该方法在真实世界同分布任务上最高提升18.7%,在分布外泛化上提升21.5%。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during finetuning, and that its quality synchronizes with both task success and out-of-distribution generalization. We further introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on out-of-distribution generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。