用多模态知识图谱提升建筑规范问答能力,支持条款级证据追溯。
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
- 构建多模态知识图谱统一表示规范文档层级与异构知识。
- 生成含17万节点、31万边的大规模知识库,支持跨条款推理。
- 结合图检索与大模型,实现可追溯的条款级精准问答,适合工程领域研究者。
建筑规范对建筑安全与可持续性至关重要。现有标准应用流程依赖关键词检索与人工跨条款解读,难以支持多条款推理、多模态知识利用及条款级证据追溯。为此,本研究提出BEST-KAG框架,通过三方面改进:1)构建多模态知识图谱(MKG),统一表征文档层级与异构标准知识;2)设计规则-大模型混合知识构建管道,从251项建筑规范中提取知识,建成包含171,652个节点、310,914条边的大型知识库(MAG);3)采用基于图检索的知识增强生成架构,实现条款级定位与可追溯问答。实验表明,BEST-KAG在专家评估及BLEU、ROUGE等指标上均优于主流大模型,最佳提升达74.01%。
原文摘要 · Abstract (English)
Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, multimodal knowledge utilization, or traceable clause-level evidence linkage. To address these limitations, this study develops a multimodal knowledge-driven framework that supports question answering on standard knowledge named BEST-KAG (Knowledge-Augmented Generation for Building Engineering STandards). The framework introduces 1) a multimodal knowledge graph (MKG) for unified representation of document hierarchy and heterogeneous standard knowledge with various connections, 2) a rule-LLM hybrid knowledge construction pipeline for scalable multimodal knowledge extraction, creating a large MAG with 251 building engineering standards, 171,652 nodes and 310,914 edges, and 3) a graph-retrieval-based knowledge-augmented generation architecture for clause-grounded and traceable question answering. Experiments demonstrate that BEST-KAG consistently outperforms multiple mainstream LLMs in terms of Expert evaluation, and metrics including BLEU, and ROUGE, with the best improvement up to 74.01% compared to the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。