arXiv:2507.13490cs.CL2025-07EMNLP被引 7

检验大模型价值观探测方法的可靠性与表达能力

Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?

  • 对比三种主流探测策略,发现其对输入扰动敏感
  • 价值观探测结果与真实行为关联弱,反映能力有限
  • 适合关注AI伦理评估可靠性的研究者参考

大型语言模型(LLM)的价值观评估对跨群体用户体验具有重要影响。然而,现有方法仍面临挑战:一方面,尽管多选题(MCQ)设置已被证明易受扰动影响,但缺乏对多种探测方法的系统性比较;另一方面,探测出的价值观在多大程度上反映模型对现实行为的偏好尚不明确。本文评估了三种广泛使用的探测策略在鲁棒性和表达性上的表现。通过改变提示和选项,发现所有方法在输入扰动下均存在显著波动。此外,我们引入两项任务,分别考察价值观是否响应人口背景差异,以及是否与模型在价值相关场景中的行为一致。结果表明,人口背景对自由文本生成影响较小,且探测值与模型对价值导向行为的偏好仅呈弱相关。本研究强调需更谨慎地审视大模型价值观探测的局限性。

原文摘要 · Abstract (English)

There has been extensive research on assessing the value orientation of Large Language Models (LLMs) as it can shape user experiences across demographic groups. However, several challenges remain. First, while the Multiple Choice Question (MCQ) setting has been shown to be vulnerable to perturbations, there is no systematic comparison of probing methods for value probing. Second, it is unclear to what extent the probed values capture in-context information and reflect models' preferences for real-world actions. In this paper, we evaluate the robustness and expressiveness of value representations across three widely used probing strategies. We use variations in prompts and options, showing that all methods exhibit large variances under input perturbations. We also introduce two tasks studying whether the values are responsive to demographic context, and how well they align with the models' behaviors in value-related scenarios. We show that the demographic context has little effect on the free-text generation, and the models' values only weakly correlate with their preference for value-based actions. Our work highlights the need for a more careful examination of LLM value probing and awareness of its limitations.

大模型评测价值观探测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。