arXiv:2510.07175cs.CLcs.LG2025-10Conference of the …被引 4

发现主流大模型在心理测评中存在数据泄露,能记住题目并精准匹配分数。

Quantifying Data Contamination in Psychometric Evaluations of LLMs

  • 提出三维度量化框架:题项记忆、评估记忆、目标分数匹配
  • 21个模型测试显示BFI-44和PVQ-40存在强数据污染
  • 模型不仅能记题,还能主动调整回答以达成指定心理分数

近期研究将心理测量问卷应用于大语言模型(LLMs),以评估价值观、人格、道德基础及黑暗特质等高层心理构念。尽管已有研究警示心理量表数据污染可能威胁评估可靠性,但尚无系统性量化方法。为此,本文提出一个框架,从三个维度系统衡量心理测评中的数据污染:(1)题项记忆,(2)评估记忆,(3)目标分数匹配。对来自主要模型家族的21个模型及四个广泛使用的心理量表(如大五人格量表BFI-44和肖像价值问卷PVQ-40)进行评估,结果表明这些常用量表存在显著污染,模型不仅记忆题目,还能主动调整回答以达到特定目标分数。

原文摘要 · Abstract (English)

Recent studies apply psychometric questionnaires to Large Language Models (LLMs) to assess high-level psychological constructs such as values, personality, moral foundations, and dark traits. Although prior work has raised concerns about possible data contamination from psychometric inventories, which may threaten the reliability of such evaluations, there has been no systematic attempt to quantify the extent of this contamination. To address this gap, we propose a framework to systematically measure data contamination in psychometric evaluations of LLMs, evaluating three aspects: (1) item memorization, (2) evaluation memorization, and (3) target score matching. Applying this framework to 21 models from major families and four widely used psychometric inventories, we provide evidence that popular inventories such as the Big Five Inventory (BFI-44) and Portrait Values Questionnaire (PVQ-40) exhibit strong contamination, where models not only memorize items but can also adjust their responses to achieve specific target scores.

大模型评测数据污染心理测量语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。