通过关键节点调度提升长时智能体系统的稳定性与复用性
Alignment in Time: Peak-Aware Orchestration for Long-Horizon Agentic Systems
- 基于行为代理检测轨迹不稳,定位峰值和结尾等关键段落进行修复
- 在固定算力预算下,显著提升任务轨迹质量与可复用概率
- 适合长期自主任务系统开发,无需修改模型权重
传统AI对齐主要关注单个模型输出,但长时程任务中的自主智能体需在整个交互轨迹中保持持续可靠性。我们提出APEMO(情感感知的峰值-结尾调制调度层),一种运行时调度机制,通过操作时间-情感信号,在固定算力预算下优化计算资源分配。APEMO不修改模型权重,而是通过行为代理检测轨迹不稳定性,将修复聚焦于关键片段,如峰值时刻和结束阶段。在多智能体仿真及基于LLM的规划-执行流程中评估表明,APEMO在轨迹级质量与可复用性上均优于结构化调度器。结果将对齐重构为时间控制问题,为长时程智能体系统提供了稳健的工程路径。
原文摘要 · Abstract (English)
Traditional AI alignment primarily focuses on individual model outputs; however, autonomous agents in long-horizon workflows require sustained reliability across entire interaction trajectories. We introduce APEMO (Affect-aware Peak-End Modulation for Orchestration), a runtime scheduling layer that optimizes computational allocation under fixed budgets by operationalizing temporal-affective signals. Instead of modifying model weights, APEMO detects trajectory instability through behavioral proxies and targets repairs at critical segments, such as peak moments and endings. Evaluation across multi-agent simulations and LLM-based planner--executor flows demonstrates that APEMO consistently enhances trajectory-level quality and reuse probability over structural orchestrators. Our results reframe alignment as a temporal control problem, offering a resilient engineering pathway for the development of long-horizon agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。