arXiv:2608.19922cs.LGcs.CY2026-08被引 1

检验纽约州预测模型与实地核查的差异,发现大量记录不符且存在系统性漏洞。

Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York

  • 用预测模型替代实地检查服务管线材质,仅靠记录比对而非成对验证。
  • 纽约市4.3%-31.9%的预测结果含铅,而实际验证仅1.5%-14.5%,模型无区分能力。
  • 7,782个1940年前建筑被误标为非铅,且数据源来自客户侧模型输出。

根据美国铅铜规则修订版,公用事业公司可用预测模型判断管线材质,无需现场检查。纽约州按地址公开所用方法。几乎无地址同时具备模型分类与物理验证,因此只能在同一家公用事业公司内进行群体比对。我们筛查了全部153个至少分类100个地址的纽约本地辖区。75个(49%)覆盖125,990个地址,占筛选总量的57%,仅记录单一值。零方差本身不构成违规:其中68个与自身验证一致或样本不足;另有7个被自家施工队推翻,6个超出采样解释范围。五个属纽约市行政区,一个为相距550公里的东罗切斯特。纽约市最大:43,215个地址基于模型记录为“已知其他”,121,779个记录为“未知”,其中1,880个已挖掘,120,692个记录为铅。模型组中这两项均为零,95%置信上限为0.0085%。其余纽约地区该方法对176,888个地址中的12.21%标记为铅或“可能为铅”,其对比存在局限性。模型组建筑更年轻,中位建造年份为1984年,而实地组为1930年,建造年代解释约三分之一差异,但固定年代后,记录分类含铅率4.3%-31.9%,物理验证为1.5%-14.5%,模型未体现任何年代差异。六个年代感知估计器显示预期含铅管线数在1,150-1,450之间。两项无需比较:7,782个地址位于1940年前建筑,且2025年归档快照显示公共侧判定源自客户侧模型输出。

原文摘要 · Abstract (English)

Under the US Lead and Copper Rule Revisions, a utility may determine a service line's material with a predictive model instead of inspecting it. New York State publishes, per address, which method was used. Almost no address carries both a model classification and a physical verification, so the check is between populations within a utility rather than paired addresses. We screen all 153 New York localities that classified at least 100 addresses this way. Seventy-five (49%), covering 125,990 addresses or 57% of those screened, record one value. Zero variance alone is not misconduct: 68 of the 75 match their own verification or have too little to test. Seven are contradicted by their own crews, six beyond any sampling explanation. Five are boroughs of New York City, which file as one system; one is East Rochester, 550 km away. New York City is the largest case: a predictive model is the recorded basis for 43,215 addresses, and on all of them the recorded material is "Known Other". The city records "Unknown" on 121,779 addresses, 1,880 already excavated, and lead on 120,692. In the model bucket both counts are zero, and the 95% upper bound on the rate is 0.0085%. Across the rest of New York the same method records lead or the hedge "Unknown but could be lead" on 12.21% of 176,888 addresses, a comparison whose weaknesses we report. The model-cleared population is newer, median year built 1984 against 1930, and construction era accounts for about a third of the gap and not the rest: holding era fixed, records-based classification finds lead at 4.3-31.9%, physical verification at 1.5-14.5%, the model in no era. Six era-aware estimators place the expected lead lines among them at 1,150-1,450. Two findings need no comparison: 7,782 of these addresses are in pre-1940 buildings, and the archived 2025 snapshot shows the public-side determination was copied from a customer-side model output.

公共安全预测模型数据审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。