用智能路由系统自动选专家,提升基础设计自动化准确率。
Investigating the Potential of Large Language Model-Based Router Multi-Agent Architectures for Foundation Design Automation: A Task Classification and Expert Selection Study
- 通过路由机制动态选择专家模型完成任务分类与设计。
- 浅基础设计准确率达95.00%,桩基设计达90.63%,优于单模型3~8个百分点。
- 适合需要高可靠性的土木工程自动化场景,可作专业辅助工具。
本研究探索基于大语言模型的路由式多智能体系统在基础设计自动化中的潜力,通过智能任务分类与专家选择实现高效计算。评估了三种方案:单智能体处理、设计-校核多智能体架构,以及路由式专家选择。实验使用DeepSeek R1、ChatGPT 4 Turbo、Grok 3和Gemini 2.5 Pro在浅基础与桩基设计场景中进行测试。路由配置分别取得95.00%与90.63%的准确率,较独立Grok 3提升8.75与3.13个百分点;相比传统代理流程提升10.0至43.75个百分点。Grok 3在无外部工具下表现最优,显示其在工程数学推理上的进步。双层分类框架能有效区分基础类型,适配相应分析方法。结果表明路由式多智能体系统是基础设计自动化的最优方案,且满足专业文档标准。鉴于土木工程的安全性要求,仍需人工监督,该系统定位为先进计算辅助工具而非全自主设计替代。
原文摘要 · Abstract (English)
This study investigates router-based multi-agent systems for automating foundation design calculations through intelligent task classification and expert selection. Three approaches were evaluated: single-agent processing, multi-agent designer-checker architecture, and router-based expert selection. Performance assessment utilized baseline models including DeepSeek R1, ChatGPT 4 Turbo, Grok 3, and Gemini 2.5 Pro across shallow foundation and pile design scenarios. The router-based configuration achieved performance scores of 95.00% for shallow foundations and 90.63% for pile design, representing improvements of 8.75 and 3.13 percentage points over standalone Grok 3 performance respectively. The system outperformed conventional agentic workflows by 10.0 to 43.75 percentage points. Grok 3 demonstrated superior standalone performance without external computational tools, indicating advances in direct LLM mathematical reasoning for engineering applications. The dual-tier classification framework successfully distinguished foundation types, enabling appropriate analytical approaches. Results establish router-based multi-agent systems as optimal for foundation design automation while maintaining professional documentation standards. Given safety-critical requirements in civil engineering, continued human oversight remains essential, positioning these systems as advanced computational assistance tools rather than autonomous design replacements in professional practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。