arXiv:2608.00473cs.CVcs.AI2026-08

检验视觉语言模型在建筑图中跨视角保持构件一致性的能力

CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

论文配图:CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings
图 1 · 摘自论文原文
  • 用锚点诊断模型是否准确识别构件并外化几何关系
  • GPT-5.5在分类任务中表现最佳,自由定位精度普遍较低
  • 适合关注建筑图纸理解与生成的AI研究者使用

建筑图纸违背多视角推理常规假设:平面和剖面是截切图,而立面是立面投影,对应构件的外观变化无法由相机运动解释。本文提出CrossProjection,一种基于锚点的诊断方法,评估视觉语言模型在异构建筑视图间保持组件身份与外化几何的能力。通过类别判断、候选选择和自由点线区域定位,测试匹配性、注册性和几何定位。在23个真实图纸集上,每模型1,954个条件测试中,GPT-5.5得分82.4%,Qwen3-VL-32B-Instruct为62.2%,GLM-4.5V为57.2%。200目标对照实验显示,在自然图纸上,GPT点/区域[email protected]为54-76%,Qwen为8-10%,GLM为14-36%;线端点[email protected]为22%、4%、0%。坐标网格部分恢复了GPT的点/区域精度,但对线无效。三位建筑训练参与者达到87.3-93.3%分类准确率和76-92%真值区域命中率,表明任务可行但非人类上限。因类别未形成同一项目匹配-注册对比,且界面控制影响多重负担,不作机制性推断。结论仅限:封闭选项或标记元素成功,并不意味着具备可靠的显式几何定位能力。对于绘图引导的CAD/BIM系统,分类正确不应视为无候选空间可靠性证据。可复用图上锚点、固定分母评分和哈希锁定伪影,建立该差距的审计追踪。

原文摘要 · Abstract (English)

Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region [email protected] is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint [email protected] is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.

建筑图像几何定位视觉语言模型图纸理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。