大模型在长短回答中价值偏好不一致,需改进评估方法
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
- 用长短回答对比测试大模型价值偏好
- 短长回答间偏好相关性弱,一致性差
- 回答具体性越强,偏好越弱,跨场景表达越强
大模型伦理风险与价值倾向的评估多依赖短文本问卷和心理测量,但实际应用中涉及长篇开放回应,导致真实场景下的价值风险与偏好研究不足。本文比较短文本反应与长文本输出中的价值偏好,改变长文本中论点数量以捕捉用户表达差异。分析五种模型(llama3-8b、gemma2-9b、mistral-7b、qwen2-7b、olmo-7b)发现:(1)短文本与不同论点数的长文本间偏好相关性弱;(2)任意两种长文本生成设置间的偏好相关性同样较弱;(3)对齐仅带来价值表达一致性的轻微提升。进一步发现,论点具体性与偏好强度呈负相关,跨场景覆盖度则与偏好强度正相关。结果表明需发展更稳健的方法以保障多样化应用中价值表达的一致性。
原文摘要 · Abstract (English)
Evaluations of LLMs' ethical risks and value inclinations often rely on short-form surveys and psychometric tests, yet real-world use involves long-form, open-ended responses -- leaving value-related risks and preferences in practical settings largely underexplored. In this work, we ask: Do value preferences inferred from short-form tests align with those expressed in long-form outputs? To address this question, we compare value preferences elicited from short-form reactions and long-form responses, varying the number of arguments in the latter to capture users' differing verbosity preferences. Analyzing five LLMs (llama3-8b, gemma2-9b, mistral-7b, qwen2-7b, and olmo-7b), we find (1) a weak correlation between value preferences inferred from short-form and long-form responses across varying argument counts, and (2) similarly weak correlation between preferences derived from any two distinct long-form generation settings. (3) Alignment yields only modest gains in the consistency of value expression. Further, we examine how long-form generation attributes relate to value preferences, finding that argument specificity negatively correlates with preference strength, while representation across scenarios shows a positive correlation. Our findings underscore the need for more robust methods to ensure consistent value expression across diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。