arXiv:2607.24435cs.CLcs.AI2026-07

通过词汇消融分析,揭示大模型在人格分类中的真实依据

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

论文配图:LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
图 1 · 摘自论文原文
  • 设计黑箱审计框架,结合词汇消融与诊断指标
  • 发现自由写作信号弱但广泛,社交媒体内容证据不足
  • 功能词和情感词仍可保留人格关联,适合模型可解释性研究

大语言模型可轻易对文本进行人格标签分配,但模型可解释性仍是开放问题。为此,我们提出LEX-EC,一个可复用的黑箱审计框架,结合出现频率与一致性诊断,以及受控词汇消融,以区分边际分布效应与在受限证据下可恢复的特质相关信号。使用该框架,我们发现不同文本类型表现显著差异:自由写作文本包含最广但较弱的信号;研究生自述中外向性关联在词汇掩码后减弱;单条社交媒体状态即使在特质平衡样本中也难以提供稳定证据,暗示内容或长度存在下限。掩码主题与人口统计信息削弱部分关联,但功能词、情感词及认知风格词汇仍可保持可观测性。语言提示能改变模型自解释,但未消除主题内容。LEX-EC联合评估分类普遍性、项级关联、校正随机性的共识度、词汇限制下的稳定性及提示敏感性,跨数据集、模型与提示,刻画特质关联随可用词汇证据的变化规律,首次将词汇方法应用于黑箱人格标注的可解释性研究。

原文摘要 · Abstract (English)

Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.

模型可解释性人格分类黑箱审计词汇消融

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。