用大模型自动评分心理问卷,准确率接近人工。
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
- 微调大模型学习人工评分标准,实现自动打分。
- 模型评分与人工评分高度一致,跨场景通用。
- 适合心理学研究与临床评估快速部署使用。
敌意归因偏差指将社交互动误解为有意攻击。模糊意图敌意问卷(AIHQ)常用于测量此偏差,包含开放题,要求参与者描述负面社交情境中的意图及应对方式。这些回答虽能反映敌意认知内容,但需人工耗时评分。本研究测试大语言模型是否可自动化评分。基于既往收集的创伤性脑损伤(TBI)患者与健康对照组(HC)完成的AIHQ数据,使用一半响应微调两个模型,另一半用于测试。结果表明,微调后模型生成的敌意归因与攻击性回应评分均与人工评分高度一致,且在模糊、故意、意外三类情景下表现稳定,并复现了TBI组与HC组在敌意归因和攻击反应上的差异。模型在独立非临床数据集上也表现出良好泛化能力。为促进应用,我们提供了本地与云端两种可访问的评分接口。研究显示,大模型可显著提升AIHQ评分效率,推动心理评估在多人群中的应用。
原文摘要 · Abstract (English)
Hostile attribution bias is the tendency to interpret social interactions as intentionally hostile. The Ambiguous Intentions Hostility Questionnaire (AIHQ) is commonly used to measure hostile attribution bias, and includes open-ended questions where participants describe the perceived intentions behind a negative social situation and how they would respond. While these questions provide insights into the contents of hostile attributions, they require time-intensive scoring by human raters. In this study, we assessed whether large language models can automate the scoring of AIHQ open-ended responses. We used a previously collected dataset in which individuals with traumatic brain injury (TBI) and healthy controls (HC) completed the AIHQ and had their open-ended responses rated by trained human raters. We used half of these responses to fine-tune the two models on human-generated ratings, and tested the fine-tuned models on the remaining half of AIHQ responses. Results showed that model-generated ratings aligned with human ratings for both attributions of hostility and aggression responses, with fine-tuned models showing higher alignment. This alignment was consistent across ambiguous, intentional, and accidental scenario types, and replicated previous findings on group differences in attributions of hostility and aggression responses between TBI and HC groups. The fine-tuned models also generalized well to an independent nonclinical dataset. To support broader adoption, we provide an accessible scoring interface that includes both local and cloud-based options. Together, our findings suggest that large language models can streamline AIHQ scoring in both research and clinical contexts, revealing their potential to facilitate psychological assessments across different populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。