研究大模型如何判断条件句是否合理,发现其依赖概率与语义,但不如人类稳定。
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
- 通过概率和语义相关性评估模型对条件句的接受度
- 模型对条件概率和语义支持均敏感,但程度因架构而异
- 大模型未必更接近人类判断,提示其推理仍不完善
条件可接受性指人们对条件句“若A,则B”在逻辑上可信程度的判断,在交流与推理中至关重要。人类判断受两个因素影响:B在给定A下的条件概率,以及A对B的语义相关性(即A是否有意义地支持B)。尽管已有研究探讨大语言模型(LLMs)对条件句的推理能力,但它们如何判断此类陈述的可接受性仍不清楚。本文系统研究了不同模型家族、规模及提示策略下LLMs的条件可接受性判断。通过线性混合效应模型与方差分析,发现模型对条件概率和语义相关性均敏感,但敏感程度随架构与提示方式变化。与人类数据对比显示,虽然模型利用了概率与语义线索,但一致性低于人类;值得注意的是,更大模型并不必然更接近人类判断。
原文摘要 · Abstract (English)
Conditional acceptability refers to how plausible a conditional statement is perceived to be. It plays an important role in communication and reasoning, as it influences how individuals interpret implications, assess arguments, and make decisions based on hypothetical scenarios. When humans evaluate how acceptable a conditional "If A, then B" is, their judgments are influenced by two main factors: the $\textit{conditional probability}$ of $B$ given $A$, and the $\textit{semantic relevance}$ of the antecedent $A$ given the consequent $B$ (i.e., whether $A$ meaningfully supports $B$). While prior work has examined how large language models (LLMs) draw inferences about conditional statements, it remains unclear how these models judge the $\textit{acceptability}$ of such statements. To address this gap, we present a comprehensive study of LLMs' conditional acceptability judgments across different model families, sizes, and prompting strategies. Using linear mixed-effects models and ANOVA tests, we find that models are sensitive to both conditional probability and semantic relevance$\unicode{x2014}$though to varying degrees depending on architecture and prompting style. A comparison with human data reveals that while LLMs incorporate probabilistic and semantic cues, they do so less consistently than humans. Notably, larger models do not necessarily align more closely with human judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。