探索大模型智能体在真实场景中的落地挑战与设计方法
Agents in the Wild: Where Research Meets Deployment
- 结合案例分析智能体在医药和金融中的成功设计模式
- 提出验证流水线、备用机制等实用故障应对策略
- 适合关注智能体部署的科研与工程人员参考
基于大语言模型的智能体系统,具备推理、规划、执行及与工具或其他智能体协作的能力,正快速从研究原型转向软件工程、科学发现和金融等领域的生产级部署。尽管学术研究聚焦于基准测试与算法创新,但实际部署带来了鲁棒性、安全性和可靠性新挑战。本教程汇聚研究者与实践者,探讨推理与规划、多智能体协同及评估方面的进展,通过制药发现与金融系统的应用案例,分析促成智能体成功的常见设计模式,并讨论针对失效模式的实用缓解策略,如验证流水线、回退机制与人工介入监督。参与者将获得领域全景认知,以及可复用的设计模式、评估清单与安全部署模板。
原文摘要 · Abstract (English)
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。