arXiv:2602.04033cs.CLcs.AI2026-02中稿 · the Workshop on Mu…综述被引 4

评估大模型价值观时,提问方式和解码策略会影响结果可信度。

On the Credibility of Evaluating LLMs using Survey Questions

  • 用思维链提示和采样解码可提升评估准确性
  • 即使平均回答接近人类,模型也可能缺乏答案间的一致性
  • 适合关注大模型价值观评估方法的科研人员

近期研究通过调整社会调查问题,以提示大语言模型(LLMs)并对比其回答与人类平均回答来评估模型的价值取向。本文指出该方法存在局限性,提示方式(直接或思维链)和解码策略(贪婪或采样)会显著影响结果,可能造成高估或低估相似性。基于跨五国三语言的世界价值调查数据,我们提出一种新指标——自相关距离,用于衡量模型在不同问题间是否保持与人类一致的回答结构关系。结果显示,即使平均一致性高,若忽略答案间的结构性关联,评估结果仍不可靠。此外,均方距离与KL散度之间的相关性较弱,说明传统指标假设答案独立并不成立。建议未来研究采用思维链提示、数十次采样解码,并结合自相关距离等多指标进行稳健评估。

原文摘要 · Abstract (English)

Recent studies evaluate the value orientation of large language models (LLMs) using adapted social surveys, typically by prompting models with survey questions and comparing their responses to average human responses. This paper identifies limitations in this methodology that, depending on the exact setup, can lead to both underestimating and overestimating the similarity of value orientation. Using the World Value Survey in three languages across five countries, we demonstrate that prompting methods (direct vs. chain-of-thought) and decoding strategies (greedy vs. sampling) significantly affect results. To assess the interaction between answers, we introduce a novel metric, self-correlation distance. This metric measures whether LLMs maintain consistent relationships between answers across different questions, as humans do. This indicates that even a high average agreement with human data, when considering LLM responses independently, does not guarantee structural alignment in responses. Additionally, we reveal a weak correlation between two common evaluation metrics, mean-squared distance and KL divergence, which assume that survey answers are independent of each other. For future research, we recommend CoT prompting, sampling-based decoding with dozens of samples, and robust analysis using multiple metrics, including self-correlation distance.

大模型评估价值观对齐提示工程评测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。