把专家经验编码成AI Agent,让非专家也能做出专业级可视化
How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering Framework
- 用分类器+RAG+规则库增强LLM,构建可自主决策的可视化AI代理
- 五类场景测试中输出质量提升206%,所有案例达到专家水平
- 适合希望快速复制专家能力的工程团队或研发机构使用
关键领域知识通常仅掌握在少数专家手中,导致组织在可扩展性和决策效率上出现瓶颈。非专家难以生成有效可视化,造成洞察不足并占用专家时间。本文通过工业案例研究,探索如何将人类领域知识捕获并嵌入到AI代理系统中。我们提出一种软件工程框架,通过在大型语言模型(LLM)中集成请求分类器、用于代码生成的检索增强生成(RAG)系统、编码化的专家规则以及统一的可视化设计原则,构建一个具备自主、反应、主动和社交行为的智能代理。在跨多个工程领域的五个场景中,由12名评估者进行测试,结果表明,该代理在输出质量上相比基线提升206%,所有情况下均达到专家评分等级,同时保持更高且更稳定的代码质量。主要贡献包括:一个自动化可视化生成的代理系统,以及一个可验证的框架,能系统性地捕获人类领域知识,并将隐性专家经验转化为可复用的AI代理,证明非专家可在专业领域实现专家级成果。
原文摘要 · Abstract (English)
Critical domain knowledge typically resides with few experts, creating organizational bottlenecks in scalability and decision-making. Non-experts struggle to create effective visualizations, leading to suboptimal insights and diverting expert time. This paper investigates how to capture and embed human domain knowledge into AI agent systems through an industrial case study. We propose a software engineering framework to capture human domain knowledge for engineering AI agents in simulation data visualization by augmenting a Large Language Model (LLM) with a request classifier, Retrieval-Augmented Generation (RAG) system for code generation, codified expert rules, and visualization design principles unified in an agent demonstrating autonomous, reactive, proactive, and social behavior. Evaluation across five scenarios spanning multiple engineering domains with 12 evaluators demonstrates 206% improvement in output quality, with our agent achieving expert-level ratings in all cases versus baseline's poor performance, while maintaining superior code quality with lower variance. Our contributions are: an automated agent-based system for visualization generation and a validated framework for systematically capturing human domain knowledge and codifying tacit expert knowledge into AI agents, demonstrating that non-experts can achieve expert-level outcomes in specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。