arXiv:2509.19314cs.CLcs.AI2025-09中稿 · publication in NCM…

用大模型重写问卷题项,降低人格测评中的社会期许偏差。

Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias

  • 用GPT-o3重写50题人格量表,使题项更中性
  • 重写后可靠性和五因素结构保留,尽责性提升
  • 适合想改进心理测评公平性的研究者参考

本研究评估大语言模型(LLM)辅助的题项中性化对人格测评中社会期许偏差的缓解效果。使用GPT-o3重写国际人格项目池大五量表(IPIP-BFM-50),203名参与者完成原始版或中性化版本,并填写马洛-克劳恩社会期许量表。结果表明,中性化版本保持了可靠的信度和五因素结构,尽责性得分上升,宜人性和开放性下降。部分题项与社会期许的相关性降低,但不一致。构型不变性成立,但度量和标度不变性未通过。研究支持人工智能中性化作为潜在但不完善的偏见缓解方法。

原文摘要 · Abstract (English)

This study evaluates item neutralization assisted by the large language model (LLM) to reduce social desirability bias in personality assessment. GPT-o3 was used to rewrite the International Personality Item Pool Big Five Measure (IPIP-BFM-50), and 203 participants completed either the original or neutralized form along with the Marlowe-Crowne Social Desirability Scale. The results showed preserved reliability and a five-factor structure, with gains in Conscientiousness and declines in Agreeableness and Openness. The correlations with social desirability decreased for several items, but inconsistently. Configural invariance held, though metric and scalar invariance failed. Findings support AI neutralization as a potential but imperfect bias-reduction method.

人格测评大模型应用社会期许偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。