通过两阶段聚类分析,让房价预测既准又看得懂。
Unveiling Location-Specific Price Drivers: A Two-Stage Cluster Analysis for Interpretable House Price Predictions
- 先按地理位置分群,再分别建模,兼顾准确与解释性。
- 相比不聚类模型,误差降低36%(GAM)和58%(LR)。
- 适合想理解房价驱动因素的购房者、卖家和分析师。
房价评估因市场地域差异而复杂。现有方法多依赖难以解释的黑箱模型,或过于简化的线性回归(LR),无法捕捉市场异质性。为此,我们提出一种两阶段聚类方法:首先基于最少位置特征对房产进行分组,再引入其他特征建模。每个簇采用线性回归(LR)或广义加性模型(GAM),在预测性能与可解释性间取得平衡。在2023年德国43,309条房产数据上验证,相较于无聚类模型,GAM的平均绝对误差降低36%,LR降低58%。图形分析揭示了不同簇间模式变化。结果表明,针对簇的洞察对提升可解释性具有关键作用,为购房者、卖家及房地产分析师提供更可靠的估值参考。
原文摘要 · Abstract (English)
House price valuation remains challenging due to localized market variations. Existing approaches often rely on black-box machine learning models, which lack interpretability, or simplistic methods like linear regression (LR), which fail to capture market heterogeneity. To address this, we propose a machine learning approach that applies two-stage clustering, first grouping properties based on minimal location-based features before incorporating additional features. Each cluster is then modeled using either LR or a generalized additive model (GAM), balancing predictive performance with interpretability. Constructing and evaluating our models on 43,309 German house property listings from 2023, we achieve a 36% improvement for the GAM and 58% for LR in mean absolute error compared to models without clustering. Additionally, graphical analyses unveil pattern shifts between clusters. These findings emphasize the importance of cluster-specific insights, enhancing interpretability and offering practical value for buyers, sellers, and real estate analysts seeking more reliable property valuations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。