用专业知识增强大模型,让非专家也能准确判断绿色雨水设施维护问题。
GSI Agent: Domain Knowledge Enhancement for Large Language Models in Green Stormwater Infrastructure
- 构建GSI领域数据集并结合检索增强生成,提升模型对专业术语的理解。
- 在真实检查场景中,文本生成质量提升三倍以上(BLEU-4从0.090到0.307)。
- 适合城市运维人员、环保工程师等需处理雨水设施问题的从业者使用。
绿色雨水基础设施(GSI)系统如透水路面、雨水花园和生物滞留设施,需持续检查与维护以保障长期性能。然而,相关领域知识分散于市政手册、法规文件和检查表中,导致非专业人士难以获取可靠指导。尽管大语言模型(LLM)具备强大推理与生成能力,但在工程场景中常缺乏专业知识,易产生错误或虚构回答,限制其在专业基建任务中的应用。本文提出GSI Agent,一种面向绿色雨水设施的领域增强型大模型框架,融合三种策略:(1)基于精选指令数据集的监督微调;(2)从市政文档构建的知识库支持的检索增强生成;(3)集成检索、上下文融合与结构化输出的代理式推理流程。同时构建了符合实际检查场景的新版GSI数据集。实验结果表明,该框架显著提升了领域性能,同时保持通用知识能力稳定:在GSI数据集上,BLEU-4从0.090提升至0.307;在通用知识数据集上,得分维持在0.304与0.305之间。结果证明,系统性地注入领域知识可有效将通用大模型适配至专业基建应用。
原文摘要 · Abstract (English)
Green Stormwater Infrastructure (GSI) systems, such as permeable pavement, rain gardens, and bioretention facilities, require continuous inspection and maintenance to ensure long-term performance. However, domain knowledge about GSI is often scattered across municipal manuals, regulatory documents, and inspection forms. As a result, non-expert users and maintenance staff may struggle to obtain reliable and actionable guidance from field observations. Although Large Language Models (LLMs) have demonstrated strong general reasoning and language generation capabilities, they often lack domain-specific knowledge and may produce inaccurate or hallucinated answers in engineering scenarios. This limitation restricts their direct application to professional infrastructure tasks. In this paper, we propose GSI Agent, a domain-enhanced LLM framework designed to improve performance in GSI-related tasks. Our approach integrates three complementary strategies: (1) supervised fine-tuning (SFT) on a curated GSI instruction dataset, (2) retrieval-augmented generation (RAG) over an internal GSI knowledge base constructed from municipal documents, and (3) an agent-based reasoning pipeline that coordinates retrieval, context integration, and structured response generation. We also construct a new GSI Dataset aligned with real-world GSI inspection and maintenance scenarios. Experimental results show that our framework significantly improves domain-specific performance while maintaining general knowledge capability. On the GSI dataset, BLEU-4 improves from 0.090 to 0.307, while performance on the common knowledge dataset remains stable (0.304 vs. 0.305). These results demonstrate that systematic domain knowledge enhancement can effectively adapt general-purpose LLMs to professional infrastructure applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。