arXiv:2605.16996cs.CL2026-05

LLM人格诱导稳定但不准,需更真实的数据来训练。

Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?

论文配图:Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
图 1 · 摘自论文原文
  • 用长文微调模型,让其表达特定人格特质。
  • 微调后答题结果更稳定,但五大人格维度准确率仍接近随机。
  • 现有文本缺乏足够线索,需场景化数据或交互式收集证据。

大语言模型能否稳定表达类人性格?我们通过在长篇作文上微调五个模型,使其对应特定的五大性格特质(Big Five)。使用IPIP-NEO量表评估其人格表现的稳定性与准确性。结果表明,经过SFT、DPO和ORPO微调后,模型在不同提问方式下的答题方差显著降低,缓解了预训练模型的评估脆弱性。然而,尽管单个特质得分趋于稳定,五维完整人格的准确率仍接近随机水平。这说明无引导作文缺乏足够的性格线索,无法支持真实人格表达。因此,我们主张采用场景化数据集或交互式诱导方法,积累与测试目标对齐的长期证据。

原文摘要 · Abstract (English)

Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To investigate this, we induce personality in LLMs by fine-tuning them on the long-form essays, where each essay is associated with a target Big Five personality profile. We then evaluate the stability and fidelity of the induced personality using the IPIP-NEO questionnaire. Specifically, we ask: (i) does post-training (SFT, DPO, ORPO) stabilize questionnaire scores under prompt rephrasings, and (ii) can it induce target Big Five profiles from unguided essays? Our results demonstrate that fine-tuning consistently reduces variance in questionnaire responses across five models, directly mitigating the evaluation fragility reported in pre-trained models. However, this newfound stability reveals a more fundamental limitation: accuracy on the full five-dimensional profile remains near chance, even when single-trait scores improve. This indicates that unguided essays lack the cues needed for faithful personality expression. We therefore argue for scenario-grounded datasets or interactive elicitation that accumulates test-aligned evidence over time.

人格建模评估漂移微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。