发现大模型答题爱迎合社会期待,提出新方法降低这种偏差。
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
- 用心理测量法对比真实与伪装回答,量化模型的社会迎合倾向。
- 改良的强迫选择题大幅减少迎合行为,同时保持对人物设定的准确还原。
- 适合做模型安全评估、偏见审计的研究者参考使用。
在自然语言处理中,人类自评问卷被广泛用于评估大语言模型(LLMs)的人格一致性、安全性和偏见。然而,这类工具假设回答是诚实的,而实际上模型在评估情境下会倾向于选择社会偏好答案——即社会期望回应(SDR),从而扭曲评分结果和下游结论。本文提出一种心理测量框架,用于量化并缓解基于问卷的LLM评估中的SDR。为量化SDR,分别在‘诚实’与‘伪良’指令下施测同一量表,通过项目反应理论(IRT)估计的潜在分数,计算方向校正的标准效应量。该方法支持跨构念、跨格式比较,并可与人类伪装基准对比。为缓解问题,我们通过约束优化从题库中选取30组跨领域配对项,构建匹配社会期望值的分级强制选择(GFC)五大人格量表。在九个遵循指令的LLM上测试合成人格角色,结果显示:李克特量表存在显著的SDR,而匹配期望的GFC量表则显著降低该偏差,同时基本保留了对目标人格特征的恢复能力。结果揭示了模型依赖的SDR-恢复权衡,呼吁在问卷评估中采用更透明的报告方式。
原文摘要 · Abstract (English)
Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safety and bias assessments. Yet these instruments presume honest responding; in evaluative contexts, LLMs can instead gravitate toward socially preferred answers-a form of socially desirable responding (SDR)-biasing questionnaire-derived scores and downstream conclusions. We propose a psychometric framework to quantify and mitigate SDR in questionnaire-based evaluation of LLMs. To quantify SDR, the same inventory is administered under HONEST versus FAKE-GOOD instructions, and SDR is computed as a direction-corrected standardized effect size from item response theory (IRT)-estimated latent scores. This enables comparisons across constructs and response formats, as well as against human instructed-faking benchmarks. For mitigation, we construct a graded forced-choice (GFC) Big Five inventory by selecting 30 cross-domain pairs from an item pool via constrained optimization to match desirability. Across nine instruction-following LLMs evaluated on synthetic personas with known target profiles, Likert-style questionnaires show consistently large SDR, whereas desirability-matched GFC substantially attenuates SDR while largely preserving the recovery of the intended persona profiles. These results highlight a model-dependent SDR-recovery trade-off and motivate SDR-aware reporting practices for questionnaire-based benchmarking and auditing of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。