arXiv:2505.04643cs.CLcs.LG2025-05被引 2

用大模型预测辅助抽样,高效估算文本中仇恨犯罪数量

Prediction-powered estimators for finite population statistics in highly imbalanced textual data: Public hate crime estimation

  • 用Transformer预测文本标签,作为辅助变量改进传统抽样估计
  • 在瑞典警方报告数据上,准确估算年度仇恨犯罪数及漏报率
  • 适合有标注数据但人力标注成本高的社会事件统计场景

在需要人工标注才能获取目标变量标签的有限文本群体中,估计总体参数具有挑战性。为此,本文将Transformer编码器神经网络的预测结果与经典的调查抽样估计方法结合,以模型预测作为辅助变量。该方法在基于瑞典警察报告的仇恨犯罪统计数据中得到验证。通过Hansen-Hurwitz估计、差值估计和分层随机抽样估计,推导出年度仇恨犯罪数量及警方漏报情况。研究表明,若有可用的标注训练数据,该方法可显著降低人工标注时间,同时提供高效率的估计结果。

原文摘要 · Abstract (English)

Estimating population parameters in finite populations of text documents can be challenging when obtaining the labels for the target variable requires manual annotation. To address this problem, we combine predictions from a transformer encoder neural network with well-established survey sampling estimators using the model predictions as an auxiliary variable. The applicability is demonstrated in Swedish hate crime statistics based on Swedish police reports. Estimates of the yearly number of hate crimes and the police's under-reporting are derived using the Hansen-Hurwitz estimator, difference estimation, and stratified random sampling estimation. We conclude that if labeled training data is available, the proposed method can provide very efficient estimates with reduced time spent on manual annotation.

文本统计抽样估计仇恨犯罪Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。