用混沌工程检测大模型多智能体系统漏洞,提升实际部署可靠性。
Assessing and Enhancing the Robustness of LLM-based Multi-Agent Systems Through Chaos Engineering
- 通过混沌工程主动注入故障,测试大模型多智能体系统的稳定性。
- 发现系统在通信失败和幻觉等问题下易出现性能崩溃。
- 适合关注AI系统鲁棒性与生产级部署的开发者与研究人员。
本研究探索了混沌工程在真实场景下增强基于大语言模型的多智能体系统(LLM-MAS)鲁棒性的应用。尽管LLM-MAS在问答、内容生成、客服自动化及决策优化等任务中具有广阔潜力,但在生产或预生产环境中仍可能面临幻觉、智能体失效及通信故障等突发问题。本文提出一种混沌工程框架,旨在主动识别这些脆弱性,评估并增强系统对故障的抵御能力,确保关键应用中的可靠运行。
原文摘要 · Abstract (English)
This study explores the application of chaos engineering to enhance the robustness of Large Language Model-Based Multi-Agent Systems (LLM-MAS) in production-like environments under real-world conditions. LLM-MAS can potentially improve a wide range of tasks, from answering questions and generating content to automating customer support and improving decision-making processes. However, LLM-MAS in production or preproduction environments can be vulnerable to emergent errors or disruptions, such as hallucinations, agent failures, and agent communication failures. This study proposes a chaos engineering framework to proactively identify such vulnerabilities in LLM-MAS, assess and build resilience against them, and ensure reliable performance in critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。