arXiv:2605.29791cs.CL2026-05被引 1

发现大模型自述与行为不一致,提出新评测框架量化这一差距。

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation

论文配图:ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation
图 1 · 摘自论文原文
  • 基于人类行为数据建立人格特质与行为的对应关系
  • 14个主流大模型均显示自我报告一致但行为偏差明显
  • 提出可插拔干预方法,提升前沿模型的行为一致性

尽管大型语言模型(LLMs)在显式自我报告中能逼真模拟人格,但在隐式行为决策上常出现偏离,暴露出显著的知识-决策差距($G_{ ext{KD}}$)。现有评测因构念效度不足、多维度混杂及分布偏差难以准确衡量此不对称性。为此,我们提出ActTraitBench,一个以人类数据为基础的评估框架,用于测量LLM的人格一致性。该框架基于实证人类数据,建立心理测量维度与行为范式的严格一一映射,并采用分位数映射进行分布校准,使LLM评分分布对齐人类基准。在14个主流大模型上的实验表明,知识-决策不对称普遍存在,更大的模型虽自我报告更一致,但行为偏差反而更强。为缓解此问题,我们进一步提出链式认知对齐(CoCA),一种即插即用的推理时干预方法,在推理能力强的前沿模型中有效提升一致性,同时暴露小型架构的能力局限。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, revealing a substantial Knowledge-Decision Gap ($G_{\text{KD}}$). Existing benchmarks struggle to measure this asymmetry due to limited construct validity, multi-dimensional entanglement, and distributional biases in LLM-based evaluation. To address these issues, we propose ActTraitBench, a human-grounded evaluation framework for measuring personality consistency in LLMs. Grounded in empirical human data, ActTraitBench establishes one-to-one mappings between psychometric facets and behavioral paradigms, and applies a Distributional Calibration via Quantile Mapping procedure to align LLM-judge score distributions with human norms. Experiments on 14 mainstream LLMs reveal a pervasive knowledge-decision asymmetry, where larger and more capable models often exhibit stronger behavioral divergence despite highly consistent self-reports. To mitigate this gap, we further introduce the Chain of Cognitive Alignment (CoCA), a plug-and-play inference-time intervention that improves alignment in reasoning-capable frontier models while exposing clear capability limitations in smaller architectures.

大模型评测人格一致性知识决策差行为验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。