用可解释模型分析美国宽带接入差距,精准定位弱势区域。
Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor Profiling

- 基于8万多普查区数据,用LightGBM预测宽带使用率差异。
- 识别出收入、教育是核心影响因素,三类典型弱势群体被划出。
- 比单纯看收入更准,能帮政府科学分配650亿基建资金。
美国通过《基础设施投资与就业法案》拨款约650亿美元用于宽带扩展,但针对这些投资的证据基础型靶向方法仍不成熟。本文提出一种可解释机器学习框架,对全国83,359个普查区进行宽带采用差异的细粒度分析。利用来自2022年美国社区调查的65个社会经济、人口与基础设施特征,训练轻量梯度提升机(LightGBM)模型,在空间五折交叉验证下达到R²=0.533,Spearman rho=0.763;州级留出交叉验证(51折)确认模型泛化能力(R²=0.525)。TreeSHAP分析表明,收入与教育为关键因子组(工程交互项吸收了其组成特征的贡献),且基于SHAP的聚类揭示三类探索性因子画像:联通良好中等水平(约4.9万区)、可负担性严重受限(约2.1万区)、农村老年型(约1.3万区)。作为筛选工具,基于机器学习的区划选择在前10%区域捕获了38.0%的总体采用差距,优于仅以收入为依据的启发式方法(35.2%,提升2.8个百分点,p<0.002,采用县-区块自举法);在后悔减少角度,该模型缩小了收入仅考虑与理想选择之间的剩余差距的19%。主要贡献在于每区的因子分解:通过SHAP识别每个区预测差距中最相关的特征组(收入/教育、偏远性、年龄),并指导差异化调查。时间稳定性检验显示,以2017年数据训练、预测2022年结果(无调查年份重叠),排名稳定性仍高(rho=0.784,注意超参数在2022年数据上调优)。
原文摘要 · Abstract (English)
The United States has allocated approximately $65 billion through the Infrastructure Investment and Jobs Act for broadband expansion, yet evidence-based methods for targeting these investments remain underdeveloped. This paper presents an explainable machine learning framework for profiling broadband adoption disparities at census-tract granularity across 83,359 tracts nationwide. Using 65 socioeconomic, demographic, and infrastructure features derived from the American Community Survey 2022, we train a LightGBM model under spatial five-fold cross-validation, achieving R^2 = 0.533 and Spearman rho = 0.763; state-held-out cross-validation (51 folds) confirms generalization (R^2 = 0.525). TreeSHAP analysis identifies income and education as the dominant factor group (with the engineered interaction term absorbing attribution from its constituent features), and SHAP-based clustering reveals three exploratory factor profiles: Well-Connected Moderate (~49K tracts), Affordability-Limited Severe (~21K tracts), and Rural-Elderly (~13K tracts). As a screening tool, ML-based tract selection captures 38.0% of the total adoption gap within the top 10% of tracts versus 35.2% for income-only heuristics (+2.8 pp, p < 0.002, county-block bootstrap); in regret-reduction terms, the model closes 19% of the remaining gap between income-only and oracle selection. The primary contribution is the per-tract factor decomposition: SHAP identifies which feature groups (income/education, rurality, age) are most strongly associated with each tract's predicted gap, and informs differentiated investigation. A temporal stability check, training on ACS 2017 and predicting ACS 2022 with zero survey-year overlap, confirms ranking stability (rho = 0.784, noting hyperparameters tuned on 2022 data).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。