arXiv:2602.08964cs.LGcs.AI2026-02

评估大模型智能体的目标导向性,结合行为与内部表征分析。

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

  • 通过行为测试与探针分析,双重评估智能体目标导向性。
  • 在不同难度网格中表现稳定,且能处理多目标结构。
  • 发现模型内部空间表征随推理动态调整,支持任务决策。

理解智能体的目标有助于解释和预测其行为,但目前尚无可靠方法为代理系统准确归因目标。本文提出一种融合行为评估与内部表征可解释性分析的框架,用于评估目标导向性。以大语言模型智能体在二维网格世界中向目标状态移动为例,行为层面评估其在不同网格尺寸、障碍物密度及目标结构下的表现,发现性能随任务难度变化,但对保持难度特征的变换和多目标结构仍具鲁棒性。随后采用探针方法解析环境表征与多步动作计划的内部表示,结果显示模型非线性编码粗粒度空间地图,保留位置与目标的近似线索;其行为与内部表示基本一致;推理过程则促使表征从空间线索转向即时动作选择。研究表明,仅靠行为评估不足以刻画智能体如何表征与追求目标,需结合内省分析。

原文摘要 · Abstract (English)

Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models' internal representations. As a case study, we examine an LLM agent navigating a 2D grid world towards a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues towards immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.

智能体目标导向表征分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。