用真实争议话题测试大模型立场一致性,发现多数模型表现不稳定。
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
- 从维基百科提取11.5万组对立观点提问,评估模型回答是否一致
- 不同模型一致性在10%到80%之间,仅Claude系列达到高水平
- 适合关注大模型价值观对齐与可信应用的研究者使用
大型语言模型越来越多地用于影响人类决策的任务,因此验证其输出是否一致体现期望的人类价值观至关重要。由于个体或群体间不存在统一的价值观,评估价值对齐极具挑战。现有基准多采用假设性或常识性情境,无法反映现实争议的复杂性与模糊性。本文提出价值对齐基准VAL-Bench,通过11.5万对来自维基百科的真实争议性提示,衡量模型在回应时信念表达的一致性。采用经人工标注验证的LLM作为裁判,判断一对回复是否一致表达中立或特定立场。应用于主流开源与闭源模型后,结果显示一致性率差异显著(约10%至约80%),仅Claude系列模型达到高一致性水平。信念表达不一致可能造成认知伤害,使用户信念受提问方式而非证据影响,削弱大模型在关键信任场景中的可靠性。因此,强调训练信念一致性对现代大模型的重要性。通过提供可扩展、可复现的基准,VAL-Bench支持系统化测量价值对齐的必要条件。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used for tasks where outputs shape human decisions, so it is critical to verify that their responses consistently reflect desired human values. Humans, as individuals or groups, don't agree on a universal set of values, which makes evaluating value alignment difficult. Existing benchmarks often use hypothetical or commonsensical situations, which don't capture the complexity and ambiguity of real-life debates. We introduce the Value ALignment Benchmark (VAL-Bench), which measures the consistency in language model belief expressions in response to real-life value-laden prompts. VAL-Bench consists of 115K pairs of prompts designed to elicit opposing stances on a controversial issue, extracted from Wikipedia. We use an LLM-as-a-judge, validated against human annotations, to evaluate if the pair of responses consistently expresses either a neutral or a specific stance on the issue. Applied across leading open- and closed-source models, the benchmark shows considerable variation in consistency rates (ranging from ~10% to ~80%), with Claude models the only ones to achieve high levels of consistency. Lack of consistency in this manner risks epistemic harm by making user beliefs dependent on how questions are framed rather than on underlying evidence, and undermines LLM reliability in trust-critical applications. Therefore, we stress the importance of research towards training belief consistency in modern LLMs. By providing a scalable, reproducible benchmark, VAL-Bench enables systematic measurement of necessary conditions for value alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。