arXiv:2608.17237cs.CVcs.AI2026-08

无需训练,用智能体直接从结构图纸中识别构件并生成可编辑模型。

Training-Free Agentic Computer Vision for Structural Component Detection in 2D Structural Framing Plans

论文配图:Training-Free Agentic Computer Vision for Structural Component Detection in 2D Structural Framing Plans
图 1 · 摘自论文原文
  • 通过几何提取与制图规则解析,无须训练即可识别构件。
  • 关键构件识别准确率超90%,尺度估计误差小于0.1%。
  • 适合建筑信息模型自动化流程,不依赖标注数据。

将结构平面图转换为可编辑的有限元模型草稿既耗时又易出错。现有构件识别系统多依赖任务特定的神经检测器,而结构工程中的语言模型通常处理文本或模型数据,而非图纸。据作者所知,本工作首次在不进行任务特定检测器训练或微调的情况下,将智能体式视觉-语言层应用于从框架图PDF中进行构件检测与模型生成。确定性阶段提取几何基元,通过尺寸比例一致性估计尺度,利用显式制图语法识别五类实体,并构建可编辑布局。智能体阶段通过确定性候选、操作特异性准入测试、变更级别审查和故障闭锁事务约束输入修正。评估使用作者生成的100张图纸基准,分为开发集(用于规则修订)与独立保留集(规则冻结后仅评估一次)。所有指标均为完整框架在保留集上的端到端结果。各图尺度估计误差均低于0.1%。柱、梁、墙、支撑、开口的召回率与精确率分别为:0.922/0.997、0.886/0.990、1.000/1.000、1.000/1.000、1.000/0.964。控制实验对三张开发图重复三次腐蚀操作,校准通过全部九次试验,构件修复在九次中有五次满足严格终态标准。因两组基准共享生成器,评估未涵盖独立绘制图纸、位图输入、分析连接性或求解器验证。

原文摘要 · Abstract (English)

Converting structural framing plans into editable finite-element model drafts is labor-intensive and susceptible to transcription errors. Existing building-component recognition systems generally depend on task-specific neural detectors, whereas language-model agents in structural engineering typically operate on text or model data rather than on drawings. To the authors' knowledge, this work is the first to apply an agentic vision-language layer to structural-component detection and model drafting from framing-plan PDFs without task-specific detector training or fine-tuning. A deterministic stage extracts geometric primitives, estimates scale by dimension-ratio consensus, recognizes five entity classes using an explicit drafting grammar, and assembles an editable layout. The agentic stage constrains typed corrections through deterministic candidates, operation-specific admission tests, change-level review, and fail-closed transactions. Evaluation used an author-generated benchmark of 100 plans, divided equally between a development half used for all rule revisions and a seed-disjoint held-out half generated after the rules were frozen and evaluated once. All scores are end-to-end results for the complete framework on the held-out half. Scale estimates were within 0.1% of the generator reference for every drawing. Recall and precision were 0.922/0.997 for columns, 0.886/0.990 for beams, 1.000/1.000 for walls, 1.000/1.000 for braces, and 1.000/0.964 for openings. A controlled study repeated two corruptions three times on three development drawings. Calibration passed all nine trials, whereas member repair satisfied every strict end-state criterion in five of nine trials. Because both benchmark halves share a generator, the evaluation does not address independently drafted plans, raster input, analytical connectivity, or solver validation.

智能体结构识别零训练图纸解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。