arXiv:2601.21926cs.RO2026-01被引 1

通过变分正则化过滤机器人操作中的冗余特征,提升视觉运动策略性能。

Information Filtering via Variational Regularization for Robot Manipulation

  • 在U-Net和DiT中引入上下文感知的高斯正则化,构建自适应信息瓶颈
  • 在三个仿真基准上显著提升任务成功率,优于基线模型并达到新SOTA
  • 方法可即插即用,实测在真实机器人部署中表现良好

基于3D视觉表征的扩散型视觉运动策略已在学习复杂机器人技能方面取得优异表现。然而,现有方法普遍采用容量过大的去噪解码器,虽能提升去噪能力,但会引入中间特征块中的冗余与噪声。我们发现,在推理时随机屏蔽主干特征或跳过DiT的中间层(不改变训练过程)可提升性能,证实了中间特征中存在任务无关噪声。为此,我们提出变分正则化(VR),一个即插即用模块,对噪声特征施加上下文条件高斯分布,并应用KL散度正则项,形成自适应信息瓶颈。在三个仿真基准RoboTwin2.0、Adroit和MetaWorld上的大量实验表明,该方法在DP3-UNet和DP3-DiT上均持续提升任务成功率,达到新SOTA。真实世界实验进一步验证了其在实际部署中的有效性。

原文摘要 · Abstract (English)

Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments.

机器人操作扩散模型信息过滤正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。