arXiv:2505.23820cs.CL2025-05ACL被引 9

LLMs在无共识任务中难以反映人类分歧,常强行站队。

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

  • 构建无共识基准,测试LLM在生成、判断、辩论中的表现
  • 作为裁判或辩手时,LLM在分歧话题上倾向选择立场
  • 适用于研究AI对人类分歧的模拟能力与对齐方法改进

随着大语言模型(LLMs)越来越多地被用于替代人类进行模型对齐,其是否能复现人类判断与偏好成为疑问,尤其在人类本身存在分歧的模糊情境中。本研究考察了LLM在三种角色中的表现:答案生成者、评判者和辩手,分别对应偏好对齐(判官)和可扩展监督(辩手)等框架,以及用户交互的典型场景。通过精心构建涵盖多种先验分歧情境的‘无共识’基准,每个案例包含两个可能立场。结果表明,尽管在开放性回答中能给出细致评估,但当作为评判或辩论者时,LLM在无共识议题上往往表现出强制立场的倾向。这凸显出在缺乏人类监督的情况下,亟需更精细的对齐方法,也说明即使人类自身意见不一,LLMs仍无法完整捕捉这种分歧。

原文摘要 · Abstract (English)

The increasing use of LLMs as substitutes for humans in ``aligning'' LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambivalent scenarios where humans disagree. This study examines the biases and limitations of LLMs in three roles: answer generator, judge, and debater. These roles loosely correspond to previously described alignment frameworks: preference alignment (judge) and scalable oversight (debater), with the answer generator reflecting the typical setting with user interactions. We develop a ``no-consensus'' benchmark by curating examples that encompass a variety of a priori ambivalent scenarios, each presenting two possible stances. Our results show that while LLMs can provide nuanced assessments when generating open-ended answers, they tend to take a stance on no-consensus topics when employed as judges or debaters. These findings underscore the necessity for more sophisticated methods for aligning LLMs without human oversight, highlighting that LLMs cannot fully capture human disagreement even on topics where humans themselves are divided.

大模型对齐人类分歧无共识任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。