为多智能体AI系统开发调试工具,支持实时修改和重置对话消息。
Interactive Debugging and Steering of Multi-Agent AI Systems
- 设计交互式界面,可浏览、编辑和重置智能体间的对话记录。
- 用户研究发现,重置消息能显著提升定位错误的效率。
- 适合开发多智能体协作系统的工程师使用。
由大语言模型驱动的自主智能体团队正被用于完成复杂任务。我们通过与五位智能体开发者访谈,识别出核心挑战:难以审查长对话以定位错误,现有工具缺乏交互式调试支持,以及配置迭代困难。基于此,我们开发了交互式多智能体调试工具AGDebugger,具备消息浏览与发送界面、历史消息编辑与重置功能,以及复杂对话历史的概览可视化。在14名参与者参与的两阶段用户研究中,我们总结出常见的智能体引导策略,并强调交互式消息重置在调试中的关键作用。研究深化了对日益重要的智能体工作流界面设计的理解。
原文摘要 · Abstract (English)
Fully autonomous teams of LLM-powered AI agents are emerging that collaborate to perform complex tasks for users. What challenges do developers face when trying to build and debug these AI agent teams? In formative interviews with five AI agent developers, we identify core challenges: difficulty reviewing long agent conversations to localize errors, lack of support in current tools for interactive debugging, and the need for tool support to iterate on agent configuration. Based on these needs, we developed an interactive multi-agent debugging tool, AGDebugger, with a UI for browsing and sending messages, the ability to edit and reset prior agent messages, and an overview visualization for navigating complex message histories. In a two-part user study with 14 participants, we identify common user strategies for steering agents and highlight the importance of interactive message resets for debugging. Our studies deepen understanding of interfaces for debugging increasingly important agentic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。