arXiv:2509.08825cs.CLcs.AI2025-09被引 45

用LLM做文本标注易出错,配置不同结果差异大,可能误判研究结论。

Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation

  • 通过微调提示词即可轻易让任意结论变显著,存在简单可操作的模型滥用风险。
  • 约31%的高精度模型和一半小模型在假设检验中得出错误结论,近似显著性时风险更高。
  • 人类标注最有效防伪,回归校正可补救但需权衡误报与漏报,适合谨慎使用LLM的研究者。

大型语言模型正快速改变社会科学研究,实现数据标注与文本分析的自动化。然而,研究人员的模型选择或提示策略等配置差异会导致输出显著波动,引发系统性偏差与随机误差,进而导致假阳性(Type I)、假阴性(Type II)、符号错误(Type S)或效应夸大(Type M)等问题。我们称此现象为LLM黑客行为。通过复现21篇已发表研究中的37个标注任务,发现仅通过少量提示改写,几乎任何结论都可被呈现为统计显著。对18种模型、2361个真实假设、总计1300万条标注的分析表明,即使遵循标准流程,意外的LLM黑客风险依然极高:状态领先模型约31%的假设得出错误结论,小型模型则达50%。尽管模型性能越高、泛化能力越强,风险越低,但高精度模型仍易受影响。效应量越大,风险越低,提示需对接近显著阈值的结果加强验证。我们评估了21种缓解技术,发现人类标注能有效防止假阳性;常用回归校正可恢复有效推断,但会牺牲类型Ⅰ与Ⅱ错误之间的平衡。论文发布实用建议以防范此类风险。

原文摘要 · Abstract (English)

Large language models are rapidly transforming social science research by enabling the automation of labor-intensive tasks like data annotation and text analysis. However, LLM outputs vary significantly depending on the implementation choices made by researchers (e.g., model selection or prompting strategy). Such variation can introduce systematic biases and random errors, which propagate to downstream analyses and cause Type I (false positive), Type II (false negative), Type S (wrong sign), or Type M (exaggerated effect) errors. We call this phenomenon where configuration choices lead to incorrect conclusions LLM hacking. We find that intentional LLM hacking is strikingly simple. By replicating 37 data annotation tasks from 21 published social science studies, we show that, with just a handful of prompt paraphrases, virtually anything can be presented as statistically significant. Beyond intentional manipulation, our analysis of 13 million labels from 18 different LLMs across 2361 realistic hypotheses shows that there is also a high risk of accidental LLM hacking, even when following standard research practices. We find incorrect conclusions in approximately 31% of hypotheses for state-of-the-art LLMs, and in half the hypotheses for smaller language models. While higher task performance and stronger general model capabilities reduce LLM hacking risk, even highly accurate models remain susceptible. The risk of LLM hacking decreases as effect sizes increase, indicating the need for more rigorous verification of LLM-based findings near significance thresholds. We analyze 21 mitigation techniques and find that human annotations provide crucial protection against false positives. Common regression estimator correction techniques can restore valid inference but trade off Type I vs. Type II errors. We publish a list of practical recommendations to prevent LLM hacking.

LLM安全文本标注研究可重复性模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。