arXiv:2609.02634cs.RO2026-09

通过聚类分析揭示视觉语言动作模型的内部表征机制

Latent Cluster Analysis for Vision-Language-Action Models

论文配图:Latent Cluster Analysis for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 构建潜空间聚类框架LAVLA,分析模型各层特征
  • 加权聚类比基线提升性能,中层特征更精细稳定
  • 提取可解释语义概念,助力机器人系统理解

视觉-语言-动作(VLA)模型在机器人领域日益广泛应用,因其能将语言与感知转化为行动,但其行为背后的内部表征仍不清晰。本文提出LAVLA框架,对当前最先进的GR00T N1.5模型进行分层分析,重点关注其动作解码器。为更好刻画动作扩散过程中的潜空间,引入基于交叉注意力的嵌入加权方法,增强相关特征并抑制低信息量特征。定量评估表明,加权聚类始终优于基线。为进一步提升可解释性,从每个聚类中提取人类可读的概念,将潜表示与语义描述关联。分析显示,潜空间聚类逐步解耦时空与运动学特征,中层表示更精细,输出层趋于稳定。LAVLA提升了语言驱动机器人系统的可解释性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a framework for latent cluster analysis of VLA models, and conduct a layer-wise study of the state-of-the-art GR00T N1.5 model, with particular focus on its action decoder. To better characterise the latent space during action diffusion, we introduce a cross-attention-based embedding-weighting method that amplifies relevant features while suppressing less informative ones. Quantitative evaluation shows that weighted clustering consistently outperforms the baseline. To improve interpretability, we extract human-interpretable concepts for each cluster, linking latent representations to semantic descriptions. Our analysis shows that latent clusters progressively disentangle spatiotemporal and kinematic features, with representations becoming more refined in the middle layers and stabilising toward the output. As such, LAVLA advances the interpretability of language-driven robotic systems.

视觉语言动作潜空间分析可解释性机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。