对比了不同输入条件下住宅能耗预测模型的精度,发现真实场景下数据受限会显著影响效果。
Full-Feature versus Limited-Input Machine Learning for Residential Energy Estimation: A Comparative Analysis of RECS and ResStock Under Realistic Input Constraints
- 用树模型在全量特征和十项易得特征下对比预测性能
- 十项输入时模型R2达0.61~0.62,物理信息缺失导致性能下降
- 针对同质化群体可提升至R2=0.85,适合精准建模场景
住宅能耗估算常需在缺乏详细围护结构、设备效率、渗透率或计量数据前提下进行。本研究基于美国全国代表性数据集RECS(调查数据)与ResStock(模拟数据),量化预测精度与输入可得性之间的权衡。首先使用全特征模型建立基准性能:在全特征下,CatBoost表现最优,ResStock R²达0.90,RECS为0.73。随后将输入限制为十项可通过住户、行政记录或气象数据获取的低负担变量,模拟真实部署条件,此时模型性能收敛至RECS R²=0.61,ResStock R²=0.62,表明算法复杂度无法弥补物理与行为信息缺失。但在更同质化的ResStock子集(单户独立住宅、天然气供暖、气候区6A、2000–2010年建造)中,简化模型准确率提升至R²=0.85,证明针对性建模对均质群体有效。结果表明,树模型可作为国家级住宅能耗数据的高保真代理,但需综合考虑特征可得性、数据来源(实测/仿真)及应用场景。
原文摘要 · Abstract (English)
Residential energy estimates are often needed before detailed envelope characteristics, equipment efficiencies, infiltration, sensor, or billing data are available. This study quantifies the trade-off between predictive accuracy and input accessibility using two nationally representative U.S. residential-energy datasets: the survey-based Residential Energy Consumption Survey (RECS) and the simulation-based ResStock dataset. Full-feature models were first used to establish dataset-specific performance benchmarks. For total-energy estimation, the models were subsequently restricted to ten low-burden variables obtainable from occupants, administrative records, or location-based weather data without an on-site energy audit. Among CatBoost, XGBoost, LightGBM, Random Forest, and Neural Networks, CatBoost consistently achieved the highest predictive performance for the full-feature analysis, reaching R2 = 0.90 for ResStock and R2 = 0.73 for RECS. When the feature set was restricted to ten homeowner-accessible inputs to simulate realistic deployment conditions, model performance converged to R2 = 0.61 for RECS and R2 = 0.62 for ResStock, showing that algorithmic complexity cannot fully compensate for missing physical and behavioral information. However, for a more homogeneous ResStock cohort consisting of single-family detached, natural-gas-heated homes in Climate Zone 6A constructed between 2000 and 2010, a reduced-input model improved accuracy to R2 = 0.85, demonstrating the value of targeted modeling for homogeneous populations. The results indicate that tree-based ensemble models can serve as high-fidelity emulators of national-scale residential energy datasets. However, careful consideration of feature availability, dataset origin (empirical vs. synthetic), and applicable use cases are also important.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。