arXiv:2608.15032cs.CLcs.AI2026-08

用AI系统自动从建筑蓝图提取材料数量,准确率超人工估算。

Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints

论文配图:Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints
图 1 · 摘自论文原文
  • 分三层构建:视觉模型提取图形元素,工具型智能体处理图纸任务,知识库支撑项目结构
  • 在10套真实住宅蓝图上达到81.6%综合得分,覆盖率达86.1%,数量精度78.8%
  • 适合建筑估价、工程自动化研究者,可复用数据与评估框架

将一组建筑蓝图转化为完整的材料数量清单,需跨图幅视觉感知、尺寸与多跳推理,并符合未明言的施工规范。我们提出Handoff-H1系统,包含三层架构:专用计算机视觉模型提取基础图形;具备图像操作与自研视觉工具(如基于CV模型的计数、检测、图纸分解)的工具型智能体;以及基于人工校准施工知识库的持久化、层级化项目基础。在建筑蓝图算量基准测试中,10套真实住宅蓝图配共识验证专家算量——共2,009项已验证条目,仅对1,348项主类材料评分。采用大模型裁判,按专业类别评估材料覆盖率与数量精确率P@25%并加权合成。同源原始PDF下,七种前沿及开源模型得分范围为35-61,独立专业估价师对照同一基准得77.6%(覆盖率65.5%,P@25% 87.9%)。Handoff-H1从原始PDF端到端运行,达81.6%(覆盖率86.1%,P@25% 78.8%),比最强前驱智能体高约20点,且在数量精度接近人类的同时实现更高覆盖率。评估框架已公开,蓝图数据与真值可申请用于研究。

原文摘要 · Abstract (English)

Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state. We present Handoff-H1, a takeoff system built from three layers: purpose-built computer-vision models that extract primitives; tool-using agents equipped with image operations and in-house visual-task tools, including CV-model-backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base. We evaluate on the Construction Blueprint Takeoff Benchmark: 10 real residential blueprint sets paired with consensus-validated expert takeoffs - 2,009 verified line items, restricted for scoring to the 1,348 primary-tier materials that drive an estimate - scored per trade by an LLM judge on material coverage and quantity Precision@25% ([email protected]) and combined into a weighted composite. Under identical scoring from the raw PDF, seven frontier and open-weight models span composites of 35-61, and independent professional estimators - scored against the same reconciled gold standard - post 77.6% (65.5% coverage, 87.9% [email protected]). Handoff-H1, working end-to-end from the raw PDF, reaches 81.6% (86.1% coverage, 78.8% [email protected]): roughly 20 points above the strongest frontier agent, and above the independent estimators by pairing near-human quantity precision with coverage they do not reach. The evaluation harness is public for the open harbor framework; the blueprint sets and ground truth are available upon request for research use.

建筑信息建模视觉推理智能估价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。