诊断视觉语言模型如何将信念信息转化为行动决策
CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models

- 通过跨轴激活引导诊断模型内部信念与行动的关联机制
- 发现模型无法将伙伴的信念信息用于下一步动作预测
- 适用于研究多智能体协作与模型决策可解释性的读者
线性探测和激活操控揭示了视觉语言模型(VLMs)在内部表征了代理的信念、知识和意图等心理状态。然而,这些表征是否以及如何被下游预测所利用仍不明确。为填补这一空白,我们提出跨轴路由诊断方法(CARD),在沿一个轴引导激活的同时,测量另一轴预测的响应。该方法应用于开放权重的VLMs在新提出的合作网格世界基准Relay Chain上,诊断出一个关键路由失败:模型未能将信念表征纳入其下一步动作预测中,导致关于合作者的重要信息被忽视。
原文摘要 · Abstract (English)
Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. However, it is unclear whether and how these representations are used by downstream predictions along these axes. To close this gap, we introduce Cross-Axis Routing Diagnostic (CARD), which steers activations along one axis while measuring the response of a different axis's prediction. Applied to open-weight VLMs on Relay Chain -- a new cooperative grid-world benchmark we propose -- we diagnose a critical routing failure: models fail to incorporate belief representations into their next action prediction, effectively leaving valuable information about their partners unused.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。