用外部验证器闭环反馈,让大模型生成合规的结构设计。
Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

- 引入物理验证器形成闭环,自动修复代码违规
- 合规率从56.8%提升至98.6%,材料用量减少5.8%
- 适合安全关键场景,开源可复现
多智能体大语言模型用于结构设计,但多数采用一次性生成且无法验证输出,难以胜任高安全性任务。本框架将外部基于物理的验证器引入闭环修复流程,耦合三层有限元验证系统与双节点循环:节点1将规范违规转为硬约束,节点2将四维质量评分转化为安全优先的软约束,检索增强的规范库使每处违规可追溯到具体条款。在44个案例、多种结构类型上,合规率由56.8%升至98.6%,综合评分从63.8增至71.4(p<0.000001),材料用量减少约5.8%。移除任一节点均导致性能下降,且在两种主干LLM间合规性无显著差异,说明效果主要来自外部验证器而非模型本身。框架、44例基准数据及全部实验脚本均已开源。
原文摘要 · Abstract (English)
Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving themill-suited to safety-critical tasks.Rather than trusting LLM self-correction,thisframework injects feedback from an external physics-based verier into a closedrepair loop.The framework couples a three-layernite-element verication systemwith a dual-node loop.Node 1 turns code violations into hard repair constraints,Node 2 turns a four-dimensional quality score into safety-rst soft constraints,and a retrieval-augmented code base makes every violation traceable to a clause.Overve structure types and 44 cases,code compliance rises from 56.8%to 98.6%and the composite score from 63.8 to 71.4(p<0.000001),using about 5.8%lessmaterial.Removing either node degrades performance,and compliance does notchange detectably across the two backbone LLMs tested,indicating that it ishere attributed to the external verier rather than the model.The framework,the 44-case benchmark and all experiment scripts are released as open source forreplicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。