检验大模型价值观评估的可靠性,发现不同测试方式结果不一致。
STONIC: A Layered Measurement Contract for LLM Value Profiling

- 设计分层测量协议,对比多种评估方式的行为一致性。
- 17个配置中仅10个保持选择与认可关系,多数偏好自身早期回答。
- 提示词本身不如决策状态能反映真实价值,适合评估可信度的研究者。
LLM价值观研究常将问卷评分、成对选择和生成文本推断的价值合并为单一画像,假设三类观察反映同一稳定偏好。STONIC在来自四家银行的5,144个情境及35个固定模型配置上检验该假设。比较孤立评分、平衡冲突下的选择、自发回答以及模型自身答案与人工替代项之间的后续选择。17个具有可用行为数据的配置中,10个在各银行间保持了认可-选择关系。所有17个配置均偏好自身早期回答(中位效应0.790),尽管选项位置变化影响每种配置的选择率。画像形态从评分到冲突选择转移最强,自发文本中减弱。对200个L3响应进行三方标注,任务局部验证语义审计:FULCRA与人类多数意见最接近,而校准后DeBERTa仍保留有用排序信息。隐藏状态比提示词更清晰编码已完成决策。因此,模型表现出可复现的行为连续性,但证据不支持跨接口的单一评分独立价值身份。
原文摘要 · Abstract (English)
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。