研究大模型对可疑巧合的敏感性,发现需显式提示才表现出人类式归纳推理。
On Language Models' Sensitivity to Suspicious Coincidences
- 通过数列和城市任务测试模型对可疑巧合的反应
- 零样本下模型未表现明显巧合敏感性,但提示后出现人类相似行为
- 提示能激发模型对假设空间的探索,提升归纳能力
人类在归纳推理时对可疑巧合敏感,倾向于选择更具体而非更泛化的假设。例如面对 {Austin, Dallas, Houston},人们更倾向认为是‘德克萨斯州城市’而非‘美国城市’。这种现象与语用推理密切相关,可作为评估系统对任务沟通目标敏感性的测试基准。本文研究语言模型(LMs)是否也表现出类似效应,聚焦两个领域:1)数列游戏(如判断4是否属于{16, 32, 2}),2)将该设置扩展至知名城市。在两种场景中,数据均兼容多个假设,研究哪个假设最符合模型行为。分析五种模型发现,其零样本行为中无强证据显示对可疑巧合的敏感性;但当通过思维链或显式提示提供假设空间后,模型开始表现出类似人类的巧合敏感效应,有时甚至与人类一致。结果表明,诱导推理行为可通过显式暴露假设空间实现。
原文摘要 · Abstract (English)
Humans are sensitive to suspicious coincidences when generalizing inductively over data, as they make assumptions as to how the data was sampled. This results in smaller, more specific hypotheses being favored over more general ones. For instance, when provided the set {Austin, Dallas, Houston}, one is more likely to think that this is sampled from "Texas Cities" over "US Cities" even though both are compatible. Suspicious coincidence is strongly connected to pragmatic reasoning, and can serve as a testbed to analyze systems on their sensitivity towards the communicative goals of the task (i.e., figuring out the true category underlying the data). In this paper, we analyze whether suspicious coincidence effects are reflected in language models' (LMs) behavior. We do so in the context of two domains: 1) the number game, where humans made judgments of whether a number (e.g., 4) fits a list of given numbers (e.g., 16, 32, 2); and 2) by extending the number game setup to prominent cities. For both domains, the data is compatible with multiple hypotheses and we study which hypothesis is most consistent with the models' behavior. On analyzing five models, we do not find strong evidence for suspicious coincidences in LMs' zero-shot behavior. However, when provided access to the hypotheses space via chain-of-thought or explicit prompting, LMs start to show an effect resembling suspicious coincidences, sometimes even showing effects consistent with humans. Our study suggests that inductive reasoning behavior in LMs can be enhanced with explicit access to the hypothesis landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。