让人类能直观监控并干预长时间运行的AI代理,提升掌控力与任务成功率。
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

- 通过可视化轨迹、手动与自动干预,实现对多任务并发代理的实时控制。
- 用户识别关键信息速度提升38%(p=0.023),小模型任务完成率最高提高34个百分点。
- 兼容主流开源与前沿代理框架,适合研究者与开发者用于调试和部署复杂代理系统。
AI代理在处理复杂、长期任务方面能力日益增强,但人机交互界面滞后于其自主性发展。为应对这一挑战,我们提出AgentGUI——一个本地部署、用户友好的图形化界面,支持对多个并发长周期代理会话的无缝观察与干预。AgentGUI具备三大功能:1)丰富的代理行为轨迹可视化;2)高效的手动与自动化干预机制;3)与开源及前沿代理框架的集成与协调。受控用户研究表明,使用AgentGUI可使用户识别代理日志中关键元素的时间减少38%(p=0.023)。初步实验显示,其自动漂移预防功能可使小型本地代理的任务完成率在0.8B至9B模型规模范围内平均提升34个百分点(每模型50次运行)。项目已公开,可通过官网(https://agent-gui-project.github.io)与开源仓库(https://github.com/eth-medical-ai-lab/agent-gui)获取,并附演示视频(https://youtube.com/watch?v=GSDyxN1gTF0)。
原文摘要 · Abstract (English)
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI's automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B--9B model ladder (N=50 runs per model). AgentGUI is publicly available through its project website (https://agent-gui-project.github.io) and open-source repository (https://github.com/eth-medical-ai-lab/agent-gui), along with a demo video (https://youtube.com/watch?v=GSDyxN1gTF0).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。