arXiv:2507.22358cs.AIcs.HC2025-07被引 48

打造可交互的智能体系统,让人类轻松参与并控制AI完成复杂任务。

Magentic-UI: Towards Human-in-the-loop Agentic Systems

  • 基于多智能体架构,支持网页浏览、代码执行等操作。
  • 六种交互机制实现低成本高效的人机协作,提升任务成功率。
  • 适合研究人机协同、安全智能体的开发者与研究人员使用。

由大语言模型驱动的AI智能体在使用外部工具完成复杂多步任务方面能力日益增强,但在计算机使用、软件开发和科研等领域仍无法达到人类水平。其日益增长的自主性及对外部世界的交互能力也带来了安全与风险问题,如行为偏离目标或遭受恶意操控。本文认为,引入人类监督的智能体系统是一条可行路径,能结合人类控制与AI效率,释放不完美系统的生产力。我们提出Magentic-UI,一个开源的网页界面,用于开发与研究人机交互。该系统基于灵活的多智能体架构,支持网页浏览、代码执行与文件操作,并可通过模型上下文协议(MCP)扩展多种工具。Magentic-UI提供了六种交互机制:共规划、共任务、多任务、动作守卫与长期记忆,以实现高效且低门槛的人类参与。我们在四个维度评估了该系统:智能体基准上的自主任务完成率、模拟用户测试的交互能力、真实用户的定性研究以及针对性的安全评估。结果表明,Magentic-UI在推动安全高效的智能体人机协作方面具有显著潜力。

原文摘要 · Abstract (English)

AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-level performance in most domains including computer use, software development, and research. Their growing autonomy and ability to interact with the outside world, also introduces safety and security risks including potentially misaligned actions and adversarial manipulation. We argue that human-in-the-loop agentic systems offer a promising path forward, combining human oversight and control with AI efficiency to unlock productivity from imperfect systems. We introduce Magentic-UI, an open-source web interface for developing and studying human-agent interaction. Built on a flexible multi-agent architecture, Magentic-UI supports web browsing, code execution, and file manipulation, and can be extended with diverse tools via Model Context Protocol (MCP). Moreover, Magentic-UI presents six interaction mechanisms for enabling effective, low-cost human involvement: co-planning, co-tasking, multi-tasking, action guards, and long-term memory. We evaluate Magentic-UI across four dimensions: autonomous task completion on agentic benchmarks, simulated user testing of its interaction capabilities, qualitative studies with real users, and targeted safety assessments. Our findings highlight Magentic-UI's potential to advance safe and efficient human-agent collaboration.

智能体人机交互安全开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。