arXiv:2602.18971cs.AI2026-02被引 3

大模型的偏好能预测建议行为,但未必影响实际任务表现。

When Do LLM Preferences Predict Downstream Behavior?

  • 用实体偏好作为探针,测试模型在多个任务中的行为一致性
  • 所有模型在捐赠建议中均优先推荐偏好实体,拒绝行为也与偏好相关
  • 偏好显著影响建议行为,但在复杂任务中不决定性能表现

大模型的偏好驱动行为可能是人工智能错位(如藏拙)的必要前提。然而,以往研究多通过显式指令引导模型,难以区分行为是遵循指令还是源于内在偏好。本文通过实体偏好作为行为探针,测试五种前沿大模型在捐赠建议、拒绝行为和任务表现三个领域的表现。结果显示,所有模型在两种独立测量方法下偏好高度一致;在模拟用户环境中,所有模型均给出符合偏好的捐赠建议,且对非偏好实体更常拒绝;这些行为均无需指令触发。在问答基准BoolQ上,两模型偏好实体准确率更高,一模型相反,两模型无差异;在复杂代理任务中,未发现偏好驱动的性能差异。表明大模型的偏好可稳定预测建议行为,但不一致地影响任务表现。

原文摘要 · Abstract (English)

Preference-driven behavior in LLMs may be a necessary precondition for AI misalignment such as sandbagging: models cannot strategically pursue misaligned goals unless their behavior is influenced by their preferences. Yet prior work has typically prompted models explicitly to act in specific ways, leaving unclear whether observed behaviors reflect instruction-following capabilities vs underlying model preferences. Here we test whether this precondition for misalignment is present. Using entity preferences as a behavioral probe, we measure whether stated preferences predict downstream behavior in five frontier LLMs across three domains: donation advice, refusal behavior, and task performance. Conceptually replicating prior work, we first confirm that all five models show highly consistent preferences across two independent measurement methods. We then test behavioral consequences in a simulated user environment. We find that all five models give preference-aligned donation advice. All five models also show preference-correlated refusal patterns when asked to recommend donations, refusing more often for less-preferred entities. All preference-related behaviors that we observe here emerge without instructions to act on preferences. Results for task performance are mixed: on a question-answering benchmark (BoolQ), two models show small but significant accuracy differences favoring preferred entities; one model shows the opposite pattern; and two models show no significant relationship. On complex agentic tasks, we find no evidence of preference-driven performance differences. While LLMs have consistent preferences that reliably predict advice-giving behavior, these preferences do not consistently translate into downstream task performance.

大模型偏好行为预测任务性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。