用大模型辅助住宅节能改造决策,效果不错但需改进准确性与上下文理解。
Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models
- 对比7个大模型在400套房屋上做节能改造推荐
- 技术目标下最高达54.5%匹配最优方案,社会经济目标受限于成本权衡
- 模型推理较清晰但缺乏深层情境理解,适合初筛与辅助决策
传统建筑节能改造决策方法普遍存在泛化能力差、可解释性低的问题,难以在多样化的居住环境中推广。随着智能社区发展,生成式AI特别是大语言模型(LLMs)可通过处理上下文信息生成面向实践者的可读建议。本文评估了七种LLMs(ChatGPT、DeepSeek、Gemini、Grok、Llama、Claude)在两个目标下的表现:最大化二氧化碳减排(技术目标)和最小化投资回收期(社会技术目标)。基于涵盖49个美国州的400套住宅数据集,从准确性、一致性、敏感性和推理能力四个维度进行评估。结果显示,未经过微调的模型在多数情况下能生成有效建议,技术目标下最高可达54.5%的推荐与最优解一致,92.8%在前五名内。社会技术目标受经济权衡和本地情境制约,表现受限。各模型间一致性较低,高性能模型反而更偏离其他模型。模型对地理位置和建筑几何敏感,但对技术类型和住户行为不敏感。多数模型表现出分步推导的工程式推理,但常简化且缺乏深层上下文感知。总体而言,大语言模型在节能改造决策中具有潜力,但在准确性、一致性和情境处理方面仍需改进才能投入实际应用。
原文摘要 · Abstract (English)
Conventional approaches to building energy retrofit decision making suffer from limited generalizability and low interpretability, hindering adoption in diverse residential contexts. With the growth of Smart and Connected Communities, generative AI, especially large language models (LLMs), may help by processing contextual information and producing practitioner readable recommendations. We evaluate seven LLMs (ChatGPT, DeepSeek, Gemini, Grok, Llama, and Claude) on residential retrofit decisions under two objectives: maximizing CO2 reduction (technical) and minimizing payback period (sociotechnical). Performance is assessed on four dimensions: accuracy, consistency, sensitivity, and reasoning, using a dataset of 400 homes across 49 US states. LLMs generate effective recommendations in many cases, reaching up to 54.5 percent top 1 match and 92.8 percent within top 5 without fine tuning. Performance is stronger for the technical objective, while sociotechnical decisions are limited by economic trade offs and local context. Agreement across models is low, and higher performing models tend to diverge from others. LLMs are sensitive to location and building geometry but less sensitive to technology and occupant behavior. Most models show step by step, engineering style reasoning, but it is often simplified and lacks deeper contextual awareness. Overall, LLMs are promising assistants for energy retrofit decision making, but improvements in accuracy, consistency, and context handling are needed for reliable practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。