arXiv:2601.01862cs.CLcs.IR2026-01被引 1

让大模型模拟人格特质,提升搜索相关性判断的准确性与可信度。

Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment

  • 用五大性格特质引导大模型生成不同行为模式,评估其对判断的影响。
  • 低尽责性模型在避免过度自信和低估方面表现最优,相关性更接近人工标注。
  • 结合人格特征与置信度可提升分类器性能,适合构建更可靠的评估系统。

近期研究显示,提示工程可使大语言模型(LLMs)模拟特定人格特质并产生相应行为。然而,对这些模拟人格如何影响关键网络搜索决策——尤其是相关性评估——仍缺乏深入理解。此外,很少有研究探讨人格模拟对置信度校准的影响,如过度自信或低估倾向。尽管心理学文献指出这些偏差具有特质相关性,例如高外向性常伴过度自信,高神经质易导致低估。为填补该空白,我们对多个主流及开源大模型进行了全面研究,通过提示使其模拟五大性格特质,在TREC DL 2019、TREC DL 2020和LLMJudge三个测试集上评估。每条查询-文档对输出相关性判断与自报告置信度。结果显示,低宜人性在相关性判断上始终更接近人工标注;低尽责性在抑制过度自信与低估方面表现最佳。同时,不同人格下相关性评分与置信度分布呈现系统性差异。基于此,我们将人格条件化得分与置信度作为特征输入随机森林分类器,在新数据集TREC DL 2021上超越了单一人格最优表现,即使训练数据有限。结果表明,人格衍生的置信度提供互补预测信号,有助于构建更可靠且贴近人类的模型评估体系。

原文摘要 · Abstract (English)

Recent studies have shown that prompting can enable large language models (LLMs) to simulate specific personality traits and produce behaviors that align with those traits. However, there is limited understanding of how these simulated personalities influence critical web search decisions, specifically relevance assessment. Moreover, few studies have examined how simulated personalities impact confidence calibration, specifically the tendencies toward overconfidence or underconfidence. This gap exists even though psychological literature suggests these biases are trait-specific, often linking high extraversion to overconfidence and high neuroticism to underconfidence. To address this gap, we conducted a comprehensive study evaluating multiple LLMs, including commercial models and open-source models, prompted to simulate Big Five personality traits. We tested these models across three test collections (TREC DL 2019, TREC DL 2020, and LLMJudge), collecting two key outputs for each query-document pair: a relevance judgment and a self-reported confidence score. The findings show that personalities such as low agreeableness consistently align more closely with human labels than the unprompted condition. Additionally, low conscientiousness performs well in balancing the suppression of both overconfidence and underconfidence. We also observe that relevance scores and confidence distributions vary systematically across different personalities. Based on the above findings, we incorporate personality-conditioned scores and confidence as features in a random forest classifier. This approach achieves performance that surpasses the best single-personality condition on a new dataset (TREC DL 2021), even with limited training data. These findings highlight that personality-derived confidence offers a complementary predictive signal, paving the way for more reliable and human-aligned LLM evaluators.

大模型评估人格模拟置信度校准搜索相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。