arXiv:2605.09692cs.AI2026-05

提出因果状态绑定评估框架,检验语言智能体是否真正受关键状态控制。

Causal state binding predicts action control in language agents

论文配图:Causal state binding predicts action control in language agents
图 1 · 摘自论文原文
  • 设计干预式测试,判断动作是否随关键状态变化而变化。
  • 在57,816条记录中,结构化智能体在记忆、决策等维度表现更优。
  • 适合关注智能体行为可解释性与可靠性评估的研究者。

自主语言智能体日益暴露其状态、记忆、计划与约束,但现有评估很少检验这些状态变量是否真正绑定到最终动作。本文提出因果状态绑定(causal state binding)评估框架,通过干预手段测量动作是否随事件特定的关键状态改变而改变,同时对无关线索保持不变。主要评测任务为隐藏目标的有限动作基准,评分方的干预目标在生成前设定且不透露给模型。在七个语料级单元共57,816条记录中,结构化智能体在推理、记忆、否决权和自我连续性响应上均优于高随机性对照组及组件移除实验。在Qwen2.5 7B、14B、32B和Mistral-7B的开源验证中,仅行动先验、无字段提示或打乱的关键上下文无法恢复结构化控制特征。诊断性有限动作探测显示,最小关键字段读出能恢复预定动作模式,而仅表面信息、仅行动先验或打乱字段控制则不能。在300个SWE-bench Lite任务和六种API模型上,引入无需人工标注的因果状态绑定组合后,约束纯净的任务到文件命中率@3的AUC从0.873提升至0.935。该验证聚焦于任务到文件定位,而非补丁生成或问题解决。结果支持一项评估原则:动作控制由事件特定的状态-动作绑定预测,而非输出熵、动作先验匹配或论证格式单独决定。

原文摘要 · Abstract (English)

Autonomous language agents increasingly expose traces, memories, plans and constraints, but existing evaluations rarely test whether these state variables are bound to final actions. We introduce causal state binding, an intervention-coupled evaluation framework that measures whether actions change with the event-specific decisive state while remaining invariant to irrelevant cues. The primary readout is a hidden-target finite-action benchmark in which scorer-side intervention targets are assigned before generation and withheld from the model-visible prompt. Across 57,816 scored records in seven corpus-level units, structured-agent conditions exceeded high-randomness controls and targeted component removals on reason, memory, veto and self-continuity responsiveness. Open-weight validation across Qwen2.5 7B, 14B and 32B plus Mistral-7B showed that action priors, no-field prompts and scrambled decisive context did not recover the structured-control signature. In diagnostic finite-action probes, the minimal decisive-field readout recovered the prescribed action pattern whereas surface-only, action-prior-only and scrambled-field controls did not. Across 300 SWE-bench Lite issue records and six API models, adding an oracle-free causal state-binding composite to a full non-CSB baseline increased constraint-clean issue-to-file hit@3 AUC from 0.873 to 0.935. This validation concerns issue-to-file localization, not patch application or SWE-bench issue resolution. These results support a measurement principle for agent evaluation: action control is predicted by event-specific state-action binding, not by output entropy, action-prior matching or rationale format alone.

智能体评估因果推理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。