arXiv:2505.24676cs.LG2025-05中稿 · COMPASS 2025被引 3

用OCR和机器学习还原30年代房产评估数据,助力研究种族财富不平等根源。

Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning

  • 结合传统计算机视觉与深度学习OCR,批量提取12,000+房产卡信息。
  • 对5万条记录进行自动标注,准确率超90%,支持大规模分析。
  • 无扫描件时可用建筑特征预测历史估值,适用于其他县市。

尽管美国1930年代住房政策对种族财富差距的影响已被广泛记录,但因历史房产评估档案多以纸质形式保存,学者难以量化其具体财务影响。我们提出一种可复制的数字化方法,应用于单个县,构建并发布首个公开数据集。基于公开的扫描文档,我们人工标注了超过12,000处房产卡片用于训练与验证。采用两阶段方法(经典计算机视觉+深度学习OCR)对额外50,000条记录进行自动标注。对于无扫描件的情况,我们构建基于建筑特征的回归模型来估算历史价值,并在其他县测试其泛化能力。这些低成本工具使学者、社区活动家和政策制定者能更有效地分析红线政策的历史影响。

原文摘要 · Abstract (English)

Despite well-documented consequences of the U.S. government's 1930s housing policies on racial wealth disparities, scholars have struggled to quantify its precise financial effects due to the inaccessibility of historical property appraisal records. Many counties still store these records in physical formats, making large-scale quantitative analysis difficult. We present an approach scholars can use to digitize historical housing assessment data, applying it to build and release a dataset for one county. Starting from publicly available scanned documents, we manually annotated property cards for over 12,000 properties to train and validate our methods. We use OCR to label data for an additional 50,000 properties, based on our two-stage approach combining classical computer vision techniques with deep learning-based OCR. For cases where OCR cannot be applied, such as when scanned documents are not available, we show how a regression model based on building feature data can estimate the historical values, and test the generalizability of this model to other counties. With these cost-effective tools, scholars, community activists, and policy makers can better analyze and understand the historical impacts of redlining.

历史数据OCR红线下量化研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。