让视觉语言动作模型的内部状态可读,提升机器人操作的可靠性与可解释性。
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
- 通过探针分析模型深层,发现物体属性与动作状态的符号化表示
- 在多个层级上识别出准确率超0.90的符号状态,验证了隐层语义结构
- 首次实现与认知架构融合,适合研究可解释智能体与机器人系统
视觉-语言-动作(VLA)模型有望成为通用机器人解决方案,将视觉和语言输入转化为机器人动作,但其黑箱特性与对环境变化的敏感性导致可靠性不足。相比之下,认知架构(CA)擅长符号推理与状态监控,却受限于预设执行路径的僵化。本文通过探查OpenVLA模型的隐藏层,揭示其在Llama骨干网络中对物体属性、关系及动作状态的符号表征,实现了与认知架构的集成,显著提升可解释性与鲁棒性。在LIBERO-spatial抓取任务上的实验表明,多数层中物体与动作状态的编码准确率均超过0.90,尽管未观察到物体状态先于动作状态编码的预期模式。所提出的DIARC-OpenVLA系统实现了实时状态监控,为更可靠、可解释的机器人操作奠定了基础。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental changes. In contrast, cognitive architectures (CA) excel in symbolic reasoning and state monitoring but are constrained by rigid predefined execution. This work bridges these approaches by probing OpenVLA's hidden layers to uncover symbolic representations of object properties, relations, and action states, enabling integration with a CA for enhanced interpretability and robustness. Through experiments on LIBERO-spatial pick-and-place tasks, we analyze the encoding of symbolic states across different layers of OpenVLA's Llama backbone. Our probing results show consistently high accuracies (> 0.90) for both object and action states across most layers, though contrary to our hypotheses, we did not observe the expected pattern of object states being encoded earlier than action states. We demonstrate an integrated DIARC-OpenVLA system that leverages these symbolic representations for real-time state monitoring, laying the foundation for more interpretable and reliable robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。