arXiv:2606.12426cs.CYcs.CL2026-06

大模型标注器存在社交倾向偏差,可能歪曲社会科学研究结论。

Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science

  • 测试三款7B模型在六项推特任务中的标注偏差,发现错误方向不一致。
  • Zephyr过度宽容,Mistral和Qwen过度纠正,三者均低估反对立场24-40个百分点。
  • 现有提示策略无法纠正偏差,甚至安全提示会加剧立场扭曲。

大型语言模型标注器在计算社会科学中应用日益广泛,但其对齐性错误是否影响研究结论尚不明确。我们评估了三款7B指令微调模型(Zephyr、Mistral-Instruct、Qwen2.5-Instruct)在六项TweetEval任务下四种提示条件(共72组实验)的表现,发现社交倾向性偏差并非单向。Zephyr呈现宽松偏差,系统性低估有害标签(攻击性语言:假良性率0.729,假警报率0.031)。Mistral与Qwen则表现出过度纠正,重复应用相同标签(Mistral仇恨言论假警报率FAR=0.604)。三者在堕胎立场任务上均出现中立偏差,将反对意见比例低估24至40个百分点,并夸大中立标签。我们测试的四种提示干预(中立、安全框架、去人格化、思维链)均未有效校正偏差;安全框架甚至加剧立场失真。令人震惊的是,尽管Zephyr的仇恨言论整体预估值恰好匹配真实率,但其类别条件误差在双向上均显著,造成偶然抵消,误导聚合验证。我们据此提出三部分分类法,结合诊断性FBR/FAR特征与轻量级黄金样本验证协议。核心警示:模型在聚合指标上看似校准,仍可能彻底颠倒研究者得出的实际结论。

原文摘要 · Abstract (English)

LLM annotators are increasingly used in computational social science (CSS), but it is unclear whether their alignment-shaped errors preserve the empirical conclusions a researcher would report. We audit three open-source 7B instruction-tuned models (Zephyr, Mistral-Instruct, Qwen2.5-Instruct) across six TweetEval tasks under four prompt conditions (72 cells) and find that social-desirability failures do not run in a single direction. Zephyr exhibits leniency bias, systematically under-applying harmful labels (offensive language: false benign rate 0.729, false alarm rate 0.031). Mistral and Qwen exhibit overcorrection, over-applying the same labels (Mistral hate-speech FAR = 0.604). All three models exhibit neutrality bias on abortion stance, underestimating opposition prevalence by 24 to 40 percentage points and inflating the neutral label. None of the four prompting interventions we test (neutral, safety framing, depersonalized, chain-of-thought) corrects these failures across models; safety framing can worsen stance distortion. Strikingly, Zephyr's hate-speech prevalence estimate matches the gold rate exactly while its class-conditional errors are large in both directions, an accidental cancellation that misleads aggregate validation. We translate these patterns into a three-part taxonomy with diagnostic FBR/FAR signatures and a lightweight gold-sample validation protocol. The headline for trustworthy CSS: a model that looks calibrated on aggregate metrics can still flip the substantive empirical conclusion a researcher would report.

大模型审计社会偏见标注偏差计算社会学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。