arXiv:2506.09627cs.CL2025-06EMNLP被引 5

对比两种去偏方法在有限样本下的表现,发现其效果受标注数量影响且存在权衡。

Benchmarking Debiasing Methods for LLM-based Parameter Estimates

  • 通过少量专家标注结合LLM,用DSL与PPI方法减少参数估计偏差
  • 大规模数据下两者偏差均低,但DSL在降偏和效率上更优
  • 方法表现因数据集而异,需关注有限样本下的去偏效率

大型语言模型(LLMs)为文本标注提供低成本且强大的方式,但其结果常与专家不一致,可能引入下游人口参数估计(如回归系数、因果效应)的偏差。为缓解此问题,研究者提出了设计型监督学习(DSL)和预测驱动推断(PPI)等去偏方法,这些方法通过结合少量昂贵的专家标注与大量LLM标注,理论上可实现有效估计。然而,在实际应用中常见的有限样本条件下,这些方法的表现尚不明确。本文贡献有二:第一,系统研究了各方法性能随专家标注数量变化的规律,揭示了模型偏差与专家样本量不足对结果的影响;第二,在多种任务中比较了DSL与PPI,发现尽管两者在大数据下均能实现低偏差,但DSL通常在偏差降低和经验效率上更优,且表现稳定性较差。结果表明,去偏方法存在偏差-方差权衡,亟需发展用于量化其在有限样本中效率的新指标。

原文摘要 · Abstract (English)

Large language models (LLMs) offer an inexpensive yet powerful way to annotate text, but are often inconsistent when compared with experts. These errors can bias downstream estimates of population parameters such as regression coefficients and causal effects. To mitigate this bias, researchers have developed debiasing methods such as Design-based Supervised Learning (DSL) and Prediction-Powered Inference (PPI), which promise valid estimation by combining LLM annotations with a limited number of expensive expert annotations. Although these methods produce consistent estimates under theoretical assumptions, it is unknown how they compare in finite samples of sizes encountered in applied research. We make two contributions. First, we study how each methods performance scales with the number of expert annotations, highlighting regimes where LLM bias or limited expert labels significantly affect results. Second, we compare DSL and PPI across a range of tasks, finding that although both achieve low bias with large datasets, DSL often outperforms PPI on bias reduction and empirical efficiency, but its performance is less consistent across datasets. Our findings indicate that there is a bias-variance tradeoff at the level of debiasing methods, calling for more research on developing metrics for quantifying their efficiency in finite samples.

LLM去偏参数估计效率评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。