arXiv:2506.00751cs.AIcs.LG2025-06被引 7

发现大模型说的和做的常不一致,影响信任与安全

Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?

  • 设计强制二选一测试题,对比模型表态与实际选择
  • 微调提示格式就能让模型立场反转,普遍存在于各模型中
  • 对需高可靠性的智能服务和自主任务尤为关键

近期大语言模型(LLMs)的发展凸显了将其行为与人类价值观对齐的重要性。一个关键但研究不足的问题是:模型宣称的偏好(基于一般原则的回答)与其在具体情境中表现出的真实偏好(通过决策推断)之间可能存在分歧。这种偏差严重影响模型的可解释性、可信度、推理透明度及伦理部署,尤其在高风险场景中。本文首次形式化定义并提出测量该偏好偏差的方法。我们构建了一个精心设计的提示数据集,包含一系列强制二选一任务,向四款主流大模型提问。通过比较模型在一般原则提示下的表态偏好与在上下文提示下的实际选择,使用KL散度等指标量化偏差。结果发现,提示格式的微小变化即可导致模型偏好发生转向,且这一现象在不同偏好类别和模型间普遍存在。这表明当前对模型决策能力的理解与控制仍严重不足。本研究对将大模型用于直接人机交互的服务至关重要,尤其在涉及道德、公平与社会责任的领域。同时,在大模型承担自主代理任务的未来场景中,识别此类偏差将成为不可或缺的能力。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) highlight the need to align their behaviors with human values. A critical, yet understudied, issue is the potential divergence between an LLM's stated preferences (its reported alignment with general principles) and its revealed preferences (inferred from decisions in contextualized scenarios). Such deviations raise fundamental concerns for the interpretability, trustworthiness, reasoning transparency, and ethical deployment of LLMs, particularly in high-stakes applications. This work formally defines and proposes a method to measure this preference deviation. We investigate how LLMs may activate different guiding principles in specific contexts, leading to choices that diverge from previously stated general principles. Our approach involves crafting a rich dataset of well-designed prompts as a series of forced binary choices and presenting them to LLMs. We compare LLM responses to general principle prompts stated preference with LLM responses to contextualized prompts revealed preference, using metrics like KL divergence to quantify the deviation. We repeat the analysis across different categories of preferences and on four mainstream LLMs and find that a minor change in prompt format can often pivot the preferred choice regardless of the preference categories and LLMs in the test. This prevalent phenomenon highlights the lack of understanding and control of the LLM decision-making competence. Our study will be crucial for integrating LLMs into services, especially those that interact directly with humans, where morality, fairness, and social responsibilities are crucial dimensions. Furthermore, identifying or being aware of such deviation will be critically important as LLMs are increasingly envisioned for autonomous agentic tasks where continuous human evaluation of all LLMs' intermediary decision-making steps is impossible.

大模型对齐行为一致性偏好偏差可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。