LLM在建筑暖通领域仍以辅助工具为主,尚无真正落地部署。
Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

- 系统梳理66篇论文,按应用与方法分类评估
- 仅4篇达试点证据,无人实现持续运行部署
- 适合做命名规范、文档支持等辅助任务
建筑自动化系统产生大量传感器数据,但因点名不统一、元数据缺失和文档分散,难以有效利用。本文系统回顾2023年至2026年3月间发表的66篇关于大语言模型(LLMs)在暖通空调(HVAC)运行中应用的同行评审论文。每篇论文按五类应用和三类方法分类,并评估其证据真实性、部署成熟度及LLM与物理控制决策的责任边界。研究集中于建筑能耗模拟(BEM,32篇),负荷预测仍样本不足无法得出子领域结论。仅有4篇达到试点级证据,无一实现持续运营部署。所有研究均未被列为可立即投入产业使用;3篇为近中期,63篇仅为科研阶段。然而,若干限定性、人机协同场景值得近期试验,包括点名标准化、基于文档的操作员支持、BEM工作流协助及围绕物理控制器的建议接口。传统机器学习(ML)、模型预测控制(MPC)、强化学习(RL)和基于本体的工具在高频控制、短时数值预测和结构化映射方面仍更成熟;自主代理操作和未经验证的用户代理仍处于研究阶段。当前证据表明,LLMs主要适合作为语义与流程层,而非自主控制器。未来工作应优先推进现场验证基准、在实际约束下的协同评估,以及具有有限延迟和可验证安全性的LLM-MPC/RL架构。
原文摘要 · Abstract (English)
Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point naming, missing metadata, and fragmented documentation obstruct their operational use. This systematic review analyses and codes 66 peer-reviewed studies on large language models (LLMs) for HVAC operations published between 2023 and March 2026. Each study is classified across five application families and three LLM method families and assessed for evidence realism, deployment readiness, and the responsibility boundary between the LLM and physical HVAC decisions. The corpus is concentrated in building energy modelling (BEM, 32 of 66 papers), while load forecasting remains too sparse for subfield-level conclusions. Only four studies reach pilot-level evidence, and none reports sustained operational deployment. No study was classified as ready-now for industry adoption; three were near-term and 63 research-only. Nevertheless, several bounded, human-in-the-loop uses merit near-term trials, including point-name normalisation, document-grounded operator support, BEM workflow assistance, and advisory interfaces around physics-based controllers. Conventional machine learning (ML), model predictive control (MPC), reinforcement learning (RL) and ontology-based tools remain more adopted for high-frequency control, short-horizon numerical forecasting, and well-posed ontology mapping, while autonomous agentic operation and unvalidated occupant proxies remain research-stage. Current evidence therefore supports LLMs primarily as semantic and workflow layers rather than autonomous HVAC controllers. Future work should prioritise field-validated benchmarks, orchestration evaluation under operational constraints, and LLM-MPC/RL architectures with bounded latency and verifiable safety properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。