用小模型+符号规则自动审阅LEED建筑认证文档,提升合规检查效率。
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts

- 构建神经符号流水线,结合文本匹配与数值校验处理认证文件。
- 40亿参数模型在纯文本验证中达67.3%准确率,优于更大模型。
- 图像会降低效果,提示策略需根据文档丰富程度动态调整。
LEED v4.1 BD+C认证仍依赖人工阅读数百页项目文档并手动应用信用条款阈值逻辑。本文探究小型本地部署语言模型能否有效筛查认证材料,以及符号组件如何协同工作。提出一种神经符号流程:将项目PDF对齐至LEED信用章节,用信用感知关键词提取证据,通过本地部署的40亿参数语言模型(gemma3:4b)验证合规性,并使用特定于LEED的数值检查器处理定量门槛。在四所大学建筑(484份PDF,153个信用级决策)上的实验表明,40亿参数模型(gemma3:4b)作为纯文本核心验证器表现最佳,准确率达67.3%,优于更大的80亿参数模型(llama3.1:8b)。确定性数值检查器纠正关键定量信用的计算错误,使EA-p2准确率从50%提升至100%,并在可靠提取条件下改善多个其他信用项。然而,完整神经符号配置整体准确率为61.6%,低于最佳纯文本基线,主要因提取失败及对定性类别的保守判断所致。系统性消融实验显示,加入低分辨率图纸(150-300 dpi)始终降低准确率,且提示有效性取决于建筑的真实通过率:文档丰富的项目适合使用评分标准提示,文档较少的项目则链式思维提示更优。在原始项目文档范围内,该流水线及其基线为LEED v4.1 BD+C合规验证提供了可复现的初始参考点。
原文摘要 · Abstract (English)
LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language models can perform meaningful screening of LEED documentation and how deterministic symbolic components should share that work. A neuro-symbolic pipeline is introduced that aligns project PDFs to LEED credit sections, retrieves evidence with credit-aware keyword signatures, verifies compliance with a locally hosted 4-billion-parameter language model, and applies a LEED-specific numeric checker to quantitative thresholds. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that a 4-billion-parameter model (gemma3:4b) is the strongest text-only core verifier, achieving 67.3% accuracy and outperforming a larger 8-billion-parameter model (llama3.1:8b) in this task. The deterministic numeric checker corrects arithmetic errors on key quantitative credits, moving EA-p2 from 50% to 100% accuracy and improving several other credits when required values are reliably extracted. At the same time, the full neuro-symbolic configuration achieves 61.6% overall accuracy, trailing the best text-only baseline due to extraction failures and conservative behavior on qualitative categories. Systematic ablations show that adding low-resolution drawing images (150-300 dpi) consistently reduces accuracy, and that prompt effectiveness depends on the building's ground-truth PASS rate: rubric prompts perform best on documentation-rich projects, while chain-of-thought prompts perform best on documentation-lean projects. Within the specific scope of LEED v4.1 BD+C compliance verification over raw project documentation, this pipeline and its baselines provide an initial reproducible reference point for both accuracy and failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。