大模型在不确定偏好下易误判,难以区分有解与无解场景。
Preference Reasoning under Indeterminacy in Large Language Models

- 从认知与结构双维度定义偏好推理中的不确定性
- 主流大模型在验证任务中仍无法正确识别无解情况
- 适合关注AI对齐与决策可靠性的研究者阅读
随着大语言模型演变为决策代理,偏好推理能力成为对齐、协作与集体智能的基础。然而,现实中的偏好推理本质上具有不确定性:信息可能不完整,有效解也可能不存在。我们提出,不确定性本身而非仅正确性,是人工智能推理的核心挑战。为此,我们从两个维度形式化该挑战:(i) 认知不确定性,源于不完整、部分或表达性偏好;(ii) 结构不确定性,源于标准社会选择理论下解的不存在性。在一系列层级任务中,我们发现当前最先进的语言模型系统性地无法区分确定与不确定实例,在验证设置中仍表现出错误的推理行为。
原文摘要 · Abstract (English)
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We formalize this challenge along two axes, (i) epistemic indeterminacy, arising from incomplete, partial, or expressive preferences, and (ii) structural indeterminacy, arising from the non-existence of solutions under standard social choice concepts. Across a hierarchy of tasks, we show that state-of-the-art language models systematically fail to distinguish between determined and undetermined instances, exhibiting miscalibrated reasoning even in verification settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。