arXiv:2606.16285cs.CLcs.LG2026-06被引 1

解决长程智能体记忆更新中的责任错乱问题,让记忆写入更精准。

HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents

论文配图:HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents
图 1 · 摘自论文原文
  • 通过事后信息评估记忆更新的局部效用,避免错误归因。
  • 在开放域任务中优于强基线,保持压缩上下文效率。
  • 适合需要可靠记忆机制的长周期决策系统。

长程智能体依赖记忆机制压缩交互历史,但优化记忆写入面临独特责任分配挑战:记忆更新可能因下游工具失败、观测噪声或推理错误被奖惩,而非其自身贡献。本文提出HiMPO框架,通过事后信息引导的记忆策略优化,实现记忆写入行为的低纠缠责任分配。首先,在相同预写状态条件下,比较更新前后可恢复的任务相关信 息,估计记忆更新的局部效用;随后,利用事后相关性作为有界回溯过滤器,当局部效用未被目标结果支持时衰减记忆信用。所得记忆特定优势仅作用于记忆标记,而轨迹级奖励优化其余行为。在基于裁判的开放域任务与客观压缩记忆问答任务中,HiMPO均优于强记忆基线与强化学习基线,同时保持压缩上下文效率。受控干预与实时回放实验进一步表明,HiMPO减少了工具错误导致的责任泄漏,使记忆信用与写入功能影响一致,并对噪声训练目标具有鲁棒性。

原文摘要 · Abstract (English)

Long-horizon agents rely on memory mechanisms to compress interaction history, but optimizing memory writing faces a distinct credit assignment challenge: a memory update may be rewarded or penalized due to downstream tool failures, noisy observations, or reasoning errors rather than its own contribution. We propose HiMPO, a Hindsight-Informed Memory Policy Optimization framework for assigning less-entangled credit to memory-writing actions in long-horizon agents. HiMPO first estimates the local utility of a memory update by comparing the task-relevant information recoverable from the previous and updated memories under the same pre-write state. It then uses hindsight relevance as a bounded retrospective filter that attenuates memory credit when local utility is not supported by the target outcome. The resulting memory-specific advantage is applied only to memory tokens, while trajectory-level rewards optimize the rest of the agent's behavior. Across judge-based open-domain tasks and objective compressive-memory QA, HiMPO improves over strong memory-based and RL-based baselines while preserving compressed-context efficiency. Controlled interventions and live replay studies further show that HiMPO reduces blame leakage from tool-induced errors, assigns memory credit that aligns with the functional impact of memory writes, and remains robust to noisy training targets.

记忆优化强化学习长程决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。