提出新方法提升记忆驱动智能体的自进化能力
Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution

- 通过差分信号与指数移动平均积累操作级证据
- 在16个设置中14个表现最佳,优化阶段耗能降75%
- 适合需要高效稳定训练的智能体系统研究者
在批量轨迹蒸馏中,同一记忆操作可能在不同批次间收到矛盾反馈。现有方法缺乏跨批次、操作级别的证据累积机制,难以区分真正有效的操作与偶然结果。本文提出边际优势累积(MAA),形式化要求为可对齐性与可比性。MAA构建跨批次可比的差分信号,通过指数移动平均累积每操作的带符号证据,并利用语义身份合并确保跨批次可追溯性。作为后处理架构,MAA在4个基准、4个目标模型的16个设置中,14个取得最佳表现,持续优于现有批量蒸馏基线,在多数场景下达到或超越在线方法性能,同时将优化阶段的令牌消耗降低约75%。
原文摘要 · Abstract (English)
In batch-style trace distillation, the same memory operation may receive contradictory feedback across different batches. Existing methods lack a cross-batch, operation-level evidence accumulation mechanism, making it impossible to distinguish stably effective operations from accidental hits. This paper formalizes the requirement as two structural conditions, alignability and comparability, and proposes Marginal Advantage Accumulation (MAA). MAA constructs differential signals to make them comparable across batches, accumulates signed evidence per operation via EMA, and ensures cross-batch traceability through semantic identity merging. As a post-processing architecture, MAA achieves the best results in 14 out of 16 settings across 4 benchmarks and 4 target models, consistently outperforming existing batch-level distillation baselines and matching or surpassing online alternatives in most settings, while reducing optimization-phase token consumption by approximately 75%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。