让智能体在任务切换时精准调整过往经验,避免失效。
Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents
- 基于上下文与轨迹的关联性,用上下文差异指导经验迁移。
- 无需训练,在表示层调整轨迹,适配新任务上下文。
- 跨多个基准测试表现优于现有方法,适合复杂任务场景。
自主智能体通过反思和迭代优化实现自我提升,复用成功任务轨迹作为上下文示例以辅助后续推理。然而,任务切换常引发上下文不匹配问题。现有方法或丢弃轨迹,或用启发式手段修改,导致显著微调成本或性能不可靠。我们揭示了上下文-轨迹相关性:上下文变化与轨迹变化高度平行。基于此,提出无训练方法 Bifrost(BrIdge contextual gap FoR imprOvised trajectory STeering),利用上下文差异精确引导已解决轨迹向目标任务自适应,缓解上下文转移引起的错位。轨迹调整在智能体隐藏状态表示层面进行,确保轨迹转换在共享空间中准确对齐目标上下文。在多种基准测试中,Bifrost持续优于现有轨迹复用与微调自改善方法,证明智能体可在显著上下文变化下有效利用过往经验。
原文摘要 · Abstract (English)
Autonomous agents excel in self-improvement through reflection and iterative refinement, which reuse successful task trajectories as in-context examples to assist subsequent reasoning. However, shifting across tasks often introduces a context mismatch. Hence, existing approaches either discard the trajectories or manipulate them using heuristics, leading to a non-negligible fine-tuning cost or unguaranteed performance. To bridge this gap, we reveal a context-trajectory correlation, where shifts of context are highly parallel with shifts of trajectory. Based on this finding, we propose BrIdge contextual gap FoR imprOvised trajectory STeering (Bifrost), a training-free method that leverages context differences to precisely guide the adaptation of previously solved trajectories towards the target task, mitigating the misalignment caused by context shifts. Our trajectory adaptation is conducted at the representation level using agent hidden states, ensuring trajectory transformation accurately aligns with the target context in a shared space. Across diverse benchmarks, Bifrost consistently outperforms existing trajectory reuse and finetuned self-improvement methods, demonstrating that agents can effectively leverage past experiences despite substantial context shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。