用大模型量化认知评估中的主观性,发现人格与人口统计信息是关键。
Modeling Subjectivity in Cognitive Appraisal with Language Models
- 通过微调和提示工程测试大模型对主观性的量化能力。
- 人格特质和人口统计信息显著影响主观判断表现。
- 为自然语言处理与认知科学交叉研究提供新思路。
随着语言模型在跨学科、以人为中心的研究中应用日益广泛,对其能力的期待也在不断演变。除了在传统任务上的优秀表现外,模型还需在涉及信心和人类(不)一致性的用户中心度量上表现出色,这些因素反映了主观偏好。尽管主观性在认知科学中扮演重要角色并得到广泛研究,但其与自然语言处理的交叉研究仍处于探索阶段。针对这一空白,我们通过全面实验与分析,探讨语言模型如何量化认知评估中的主观性,采用微调模型与基于提示的大语言模型(LLMs)。定量与定性结果表明,人格特质和人口统计信息对主观性测量至关重要,而现有事后校准方法往往无法达到满意性能。深入分析为未来自然语言处理与认知科学的交叉研究提供了宝贵洞见。
原文摘要 · Abstract (English)
As the utilization of language models in interdisciplinary, human-centered studies grow, expectations of their capabilities continue to evolve. Beyond excelling at conventional tasks, models are now expected to perform well on user-centric measurements involving confidence and human (dis)agreement-factors that reflect subjective preferences. While modeling subjectivity plays an essential role in cognitive science and has been extensively studied, its investigation at the intersection with NLP remains under-explored. In light of this gap, we explore how language models can quantify subjectivity in cognitive appraisal by conducting comprehensive experiments and analyses with both fine-tuned models and prompt-based large language models (LLMs). Our quantitative and qualitative results demonstrate that personality traits and demographic information are critical for measuring subjectivity, yet existing post-hoc calibration methods often fail to achieve satisfactory performance. Furthermore, our in-depth analysis provides valuable insights to guide future research at the intersection of NLP and cognitive science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。