测试大模型对概率性表达的逻辑推理能力,发现多数模型存在答案偏见。
Benchmarking LLM Competence on Logical Inference over Probability Operators

- 构建14,320条生成式提示,涵盖15种推理模板,测试概率算子推理
- 29个模型中仅9个超越随机水平,多数存在非逻辑的Yes/No偏好
- 在问题形式、动词、姓名性别和来源上均发现系统性偏差,适合评估模型可信度
不确定性表达与推理在自然语言中普遍存在,对医学、法律等高风险领域至关重要。尽管大模型在逻辑推理任务上被广泛评估,但难以区分真正的符号推理与表面模式匹配。本文提出一个针对概率算子的推理基准,包含14,320条程序生成的英语提示,覆盖十五种推理模板,系统变化问题形式、否定策略和表层内容。评估29个模型后发现,大多数模型的答案偏好与其逻辑形式无关,表现出对“是”或“否”的系统性倾向。我们以“准确率较低者”作为能力下限:即模型在“应答是”和“应答否”样本中的最低准确率。仅有9个模型超过随机水平。此外,我们测试了问题形式、动词短语、姓名性别及来源的变化,发现所有维度均存在偏差。
原文摘要 · Abstract (English)
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。