arXiv:2410.05063cs.LGcs.CV2024-10ICLR被引 12

让视觉表征按控制需求聚类,提升少样本强化学习性能

Control-oriented Clustering of Visual Latent Representation

  • 基于神经坍缩思想,使视觉特征按控制目标聚类
  • 少样本训练下测试性能提升10%至35%
  • 适合数据稀缺的视觉控制任务研究者

我们首次研究基于行为克隆学习的图像控制流程中,视觉表征空间的几何特性。受图像分类中神经坍缩(NC)现象启发,我们实证发现:在离散控制任务(如Lunar Lander)中,视觉表征按自然动作标签聚类;在连续控制任务(如Planar Pushing和Block Stacking)中,则按‘控制导向’类别聚类,该类别基于输入中物体与目标的相对位姿,或输出中专家动作诱导的物体相对位姿,每类对应一个相对位姿象限(REPO)。进一步地,我们利用此聚类规律作为算法工具,通过在视觉编码器上施加NC正则化,预训练以促进控制导向聚类。令人惊讶的是,经此预训练的编码器在端到端微调后,测试性能提升10%至35%。真实世界平面推物实验验证了该预训练策略的优势。

原文摘要 · Abstract (English)

We initiate a study of the geometry of the visual representation space -- the information channel from the vision encoder to the action decoder -- in an image-based control pipeline learned from behavior cloning. Inspired by the phenomenon of neural collapse (NC) in image classification (arXiv:2008.08186), we empirically demonstrate the prevalent emergence of a similar law of clustering in the visual representation space. Specifically, in discrete image-based control (e.g., Lunar Lander), the visual representations cluster according to the natural discrete action labels; in continuous image-based control (e.g., Planar Pushing and Block Stacking), the clustering emerges according to "control-oriented" classes that are based on (a) the relative pose between the object and the target in the input or (b) the relative pose of the object induced by expert actions in the output. Each of the classes corresponds to one relative pose orthant (REPO). Beyond empirical observation, we show such a law of clustering can be leveraged as an algorithmic tool to improve test-time performance when training a policy with limited expert demonstrations. Particularly, we pretrain the vision encoder using NC as a regularization to encourage control-oriented clustering of the visual features. Surprisingly, such an NC-pretrained vision encoder, when finetuned end-to-end with the action decoder, boosts the test-time performance by 10% to 35%. Real-world vision-based planar pushing experiments confirmed the surprising advantage of control-oriented visual representation pretraining.

视觉控制聚类学习少样本强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。