arXiv:2608.17247cs.AI2026-08

测试显式状态提示能否提升个性化代理的记忆决策能力

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

  • 构建受控反事实数据集,分离提示与状态输出的影响
  • 显式状态提示未显著提升大模型记忆策略分类准确率
  • 结果表明当前方法难以真实反映内部记忆决策机制

个性化智能体需判断检索到的用户记忆是否应使用、忽略、更新或查询。本文通过该场景建立结构化中间输出的实证审计协议:先检测数据集捷径,再隔离捆绑提示变化,验证中间标签是否与答案相关,测试分解后的语义证据,并审计供应商执行失败。480个合成开发样本曾显示状态结构提示有显著收益,但TF-IDF分析揭示词汇可分性高,无独立的‘忽略’案例。因此构建含40组四类匹配对的160例冻结反事实集,采用规则推导参考策略。在该集上,展示四种状态定义可提升准确率,但单独显式状态输出字段对Llama-3.3-70B无显著改善,对GPT-OSS-120B仅带来边际且不显著增益。提供与基准相关的状态标签虽改变预测,但因标签与策略存在确定性映射,仅为标签条件诊断,非内部机制忠实体现。家族级与种子稳定性分析显示,个体准确率夸大了反事实一致性:完整四类家族成功极为罕见。探索性后续任务要求分解语义证据亦未能提升终点路由性能;对应GPT-OSS条件因供应商端请求校验不可用。本研究仅评估策略分类,未涉及下游响应、工具操作或记忆库更新。

原文摘要 · Abstract (English)

Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associated, test decomposed semantic evidence, and audit provider-level execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics showed lexical separability and no positive standalone Ignore cases. We therefore construct a frozen 160-example controlled counterfactual set with 40 matched four-way families and rule-derived reference policies. On this set, exposing the four state definitions improves accuracy, but an isolated explicit state-output field does not significantly improve policy accuracy for Llama-3.3-70B and gives only a marginal, non-significant gain for GPT-OSS-120B. Supplying benchmark-associated state labels shifts policy predictions, but because those labels deterministically map to policies, this is a label-conditioning diagnostic rather than evidence of a faithful internal mechanism. Family-level and seed-stability analyses further show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare. An exploratory follow-up that elicits decomposed semantic evidence also fails to improve routing for the cleanly evaluated endpoint; the corresponding GPT-OSS condition was unavailable because of provider-side request validation. We evaluate policy classification only, not downstream responses, tool actions, or memory-store mutation.

记忆决策大模型评估反事实分析提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。