arXiv:2607.14896cs.SEcs.AI2026-07

让大模型在结构工程中生成可执行的工作流,并通过严格验证确保每一步都正确。

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

论文配图:StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
图 1 · 摘自论文原文
  • 用受控工具和共享状态让大模型按工程流程逐步完成任务
  • 9种配置中自动流程仅22%全程成功,但结构化验证提升至82.9%
  • 适合关注工程自动化与可信大模型应用的研究者

解决结构工程请求不仅需要单一答案,还需一系列相互依赖的成果:需求解析、可计算模型、验证记录、求解器输出、适用工程校核及最终报告。以问答或脚本生成为中心的评估可能奖励流畅输出,即使底层工作流不完整、不一致或不可执行。我们提出StructureClaw,一个以成果为中心的工作台,其中大模型代理通过受控工程技能、类型化工具、共享成果状态和本地分析后端协作;并配套StructureClaw-Bench,一个包含150个受控场景的可执行基准,覆盖标准流程、交互鲁棒性及多模态结构模型重建。其可分析的标准与多模态案例要求严格的结构模型一对一匹配及数值响应与选定分析引擎的冻结参考结果一致;交互案例则需正向澄清或恢复证据,并在适当情况下安全拒绝执行。只有所有必需断言均通过才算成功。在九种文本代理配置中,仅通用执行在保留结果中通过模型-成果检查达87.0%,但全程成功率仅22.0%;而自动StructureClaw达到82.9%。交互与多模态评估进一步揭示语义状态一致性和可执行模型重构是主要瓶颈。代码与基准已开源于https://github.com/structureclaw/structureclaw。

原文摘要 · Abstract (English)

Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, and a final report. Evaluations centered on question answering or script generation may therefore reward fluent outputs even when the underlying workflow is incomplete, inconsistent, or non-executable. We present StructureClaw, an artifact-centered workbench in which LLM agents operate through governed engineering skills, typed tools, shared artifact state, and local analysis backends, together with StructureClaw-Bench, an executable benchmark of 150 controlled scenarios spanning standard workflows, interactive robustness, and multimodal structural-model reconstruction. Its analyzable standard and multimodal cases require both strict one-to-one structural-model matching and numerical-response agreement with frozen reference responses from the selected analysis engine; interactive cases instead require positive clarification or recovery evidence together with safe non-execution when appropriate. A trial succeeds only when every fixture-required assertion passes. Across nine text-agent configurations, generic-only execution passed the model-artifact check in 87.0% of retained outcomes but achieved only 22.0% E2E Success, whereas automatic StructureClaw reached 82.9%. Interactive and multimodal evaluations further identify semantic state consistency and executable model reconstruction as the dominant remaining bottlenecks. The code and benchmark are available at https://github.com/structureclaw/structureclaw.

大模型工程自动化可执行结构工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。