arXiv:2606.07614cs.LGstat.AP2026-06

用机器学习精简尼日利亚调查数据,仍能准确测度贫困与不平等。

Measuring Poverty and Inequality with Reduced Data: A Machine Learning Approach Using Nigerian Household Data

论文配图:Measuring Poverty and Inequality with Reduced Data: A Machine Learning Approach Using Nigerian Household Data
图 1 · 摘自论文原文
  • 用随机森林筛选关键收入消费项,仅靠少数变量即可分类
  • 季节性消费预测贫困准确率达80%,年均消费达60%-65%
  • 劳动收入可有效反映不平等位置,助减少调查成本

可靠测量低收入和中等收入国家的收入与消费对监测贫困与不平等至关重要,但完整家庭调查成本高且难以定期实施。本文探讨简化调查工具是否能保留关键分配信息。利用2018/19年尼日利亚综合家庭调查-面板数据,采用随机森林递归特征消除(RF-RFE)方法,识别出最能区分福利分布中个体的收入来源、消费类别及家庭特征。分析聚焦三个结果:贫困状态、五分位排名位置以及相对于基尼不平等线的位置。调查涵盖种植后与收获后两个季节,可评估不同季节条件下的表现。结果显示,RF-RFE在少量预测变量下实现高分类精度:消费方面,贫困状态和不平等线位置可通过少量支出类别准确预测;五分位分类在季节性消费中准确率达约80%,年均消费从单次季节访问预测则为60%–65%。收入方面,贫困状态使用五个预测变量可达约90%准确率,不平等线位置主要由劳动收入捕捉。研究结果表明,机器学习方法有助于优化调查设计,降低数据需求,同时保留测度与监测贫困与不平等所需的关键分配信息。

原文摘要 · Abstract (English)

Reliable measurement of income and consumption is essential for monitoring poverty and inequality in low- and middle-income countries, yet full household surveys are costly and difficult to implement regularly. This paper examines whether reduced survey instruments can preserve key distributional information. We apply Random Forest Recursive Feature Elimination (RF-RFE) to the 2018/19 Nigeria General Household Survey-Panel to identify the income sources, consumption categories and household characteristics that best classify individuals within the welfare distribution. The analysis focuses on three outcomes: poverty status, location in the quintile distribution and position relative to the Gini-based inequality line. The survey's post-planting and post-harvest periods allow us to assess performance under different seasonal contexts. Results show that RF-RFE achieves strong classification accuracy with few predictors. For consumption, poverty status and inequality-line position are accurately predicted using a small set of expenditure categories, while quintile classification reaches about 80 percent accuracy for seasonal consumption and 60--65 percent for annual consumption predicted from a single seasonal visit. For income, poverty status reaches around 90 percent accuracy with five predictors, and inequality-line position is largely captured by labour earnings. The findings suggest that machine-learning methods can help improve survey design and reduce data requirements while retaining much of the distributional information needed to measure and monitor poverty and inequality.

贫困测量机器学习调查优化不平等

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。