arXiv:2607.20270cs.CL2026-07

研究大模型在价值观识别中的混淆模式,发现其常误判相邻价值。

Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study

论文配图:Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study
图 1 · 摘自论文原文
  • 基于舒瓦茨十类价值观设计1000条俄语情境文本,评估模型识别能力。
  • 模型准确率仅68.3%(Top-1),近半错误源于相邻价值混淆。
  • 揭示多个方向性混淆模式,为价值观评估提供新分析维度。

大型语言模型在价值观评估中日益重要,但前提是其能识别具体情境中的价值表达。本文通过舒瓦茨十类基本价值观的精确识别任务,对21个指令微调的LLM进行评估。实验使用1000条俄语情境文本,涵盖十类价值观且每条由两名独立标注者标注。在固定排序响应协议下,20个输出可靠的模型构成语义面板。整体Top-1准确率为0.683,Top-3为0.892,表明模型虽常定位到正确动机区域,但对相近选项排名不稳定。相邻价值导致50.9%的语义错误,远高于检查点特异性的24.4%基准。八种定向混淆在不同检查点和人工验证子集中反复出现,其中部分具有强不对称性,如普遍主义→利他主义、传统→遵从、安全→权力;而刺激→享乐形成双向边界。混淆严重程度因检查点而异,可能扭曲高阶价值观画像。结果呼吁结合精确度、排序恢复与定向误差分析的价值观识别评估体系。

原文摘要 · Abstract (English)

Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our evaluation set contains 1,000 Russian situational texts, balanced across the ten values and independently labeled by two human annotators per item. We evaluate 21 instruction-tuned LLM runs under a fixed ranked-response protocol; 20 runs with reliable outputs form the semantic panel. Pooled Acc@1 is 0.683 and Acc@3 is 0.892, showing that models often locate the correct motivational region while ranking close alternatives unstably. Adjacent values account for 50.9% of semantic errors, compared with 24.4% under a checkpoint-specific null. Eight directed confusions recur across checkpoints and human-confirmed subsets. Several are strongly asymmetric, including Universalism to Benevolence, Tradition to Conformity, and Security to Power, whereas Stimulation-Hedonism forms a bidirectional boundary. Their severity is checkpoint-specific and can bias higher-order value profiles. The results motivate value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.

价值观识别大模型评测心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。