arXiv:2510.22170cs.AI2025-10被引 2

用情境判断测试评估大模型行为稳定性,发现其有可测量的深层倾向。

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

  • 通过情景判断题和多维项目反应理论,量化模型行为倾向
  • 同一角色设定下行为稳定,且能预测真实表现基准分数
  • 适合研究大模型行为一致性与可信度的学者使用

人物设定广泛用于引导大语言模型(LLM)行为,但尚不清楚其是否引发稳定的性格结构或仅造成表面变化。本文提出一个框架,利用情境判断测试(SJTs)、多维项目反应理论(MIRT)和结构化合成人物设定,将模型响应视为潜在行为变量的观测值。在大规模的SJTs和人物设定数据集上,我们发现人物条件化的行为在多次运行中保持稳定,潜在特质得分可有效预测外部基准表现(如TruthfulQA、EmoBench),且MIRT揭示出一致的潜在结构。通过人工标注、基准评估和内部一致性分析验证了结果可靠性。这些特质并非人类人格,而是模型在不同情境中表现出的稳定行为倾向。结果表明,基于场景的心理测量评估比传统自评方法更可靠地衡量LLM行为,相关数据集已公开以支持后续研究。

原文摘要 · Abstract (English)

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies using situational judgment tests (SJTs), multidimensional item response theory (MIRT), and structured synthetic personas, treating responses as observations of latent behavioral variables. Across large-scale SJT and persona datasets, we find that persona-conditioned behaviors are stable across runs, latent trait scores predict external benchmarks (e.g., TruthfulQA, EmoBench), and MIRT reveals consistent latent structure. We validate these results through human annotation, benchmark evaluation, and internal consistency analyses. We interpret these traits not as human personality, but as stable behavioral tendencies expressed across contexts. Our results show that scenario-based psychometric evaluation provides a more reliable alternative to classical self-report approaches for assessing LLM behavior, and we release datasets to support further study.

心理测量大模型评估行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。