SR4-Fit算法精准预测美国国会选举,兼具高准确率与可解释性。
SR4-Fit: An Interpretable and Informative Classification Algorithm Applied to Prediction of U.S. House of Representatives Elections
- 基于人口数据构建稀疏松弛正则化规则模型,融合规则可解释性与强预测力。
- 在国会选举预测中准确率超越黑箱模型和传统规则模型,稳定生成可读规则集。
- 适用于需透明决策的选举分析、政策评估等场景,尤其适合关注机制解释的研究者。
机器学习对可解释性模型的需求日益增长,但高性能模型多为黑箱,而传统规则模型如RuleFit虽易理解却预测能力弱且不稳定。为此,我们提出稀疏松弛正则化回归规则拟合(SR4-Fit),一种新型可解释分类算法,在保持优异分类性能的同时解决上述缺陷。利用美国人口普查局的社区调查数据,我们证明该算法能以前所未有的准确性和可解释性预测众议院选举政党结果。结果显示,多数党仍是最强预测因子,但SR4-Fit揭示了此前无法在随机森林等黑箱模型中解读的人口特征组合对结果的影响。相比黑箱模型和RuleFit等可解释算法,SR4-Fit在准确性、简洁性与鲁棒性方面全面领先,生成稳定且可读的规则集,打破了可解释性与预测能力之间的传统权衡。为进一步验证其表现,我们在乳腺癌、Ecoli、page blocks、Pima Indians、vehicle和yeast六个公开分类数据集上测试,结果一致优异。
原文摘要 · Abstract (English)
The growth of machine learning demands interpretable models for critical applications, yet most high-performing models are ``black-box'' systems that obscure input-output relationships, while traditional rule-based algorithms like RuleFit suffer from a lack of predictive power and instability despite their simplicity. This motivated our development of Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), a novel interpretable classification algorithm that addresses these limitations while maintaining superior classification performance. Using demographic characteristics of U.S. congressional districts from the Census Bureau's American Community Survey, we demonstrate that SR4-Fit can predict House election party outcomes with unprecedented accuracy and interpretability. Our results show that while the majority party remains the strongest predictor, SR4-Fit has revealed intrinsic combinations of demographic factors that affect prediction outcomes that were unable to be interpreted in black-box algorithms such as random forests. The SR4-Fit algorithm surpasses both black-box models and existing interpretable rule-based algorithms such as RuleFit with respect to accuracy, simplicity, and robustness, generating stable and interpretable rule sets while maintaining superior predictive performance, thus addressing the traditional trade-off between model interpretability and predictive capability in electoral forecasting. To further validate SR4-Fit's performance, we also apply it to six additional publicly available classification datasets, like the breast cancer, Ecoli, page blocks, Pima Indians, vehicle, and yeast datasets, and find similar results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。