测试大模型类比推理的鲁棒性,发现其远不如人类稳定。
Evaluating the Robustness of Analogical Reasoning in Large Language Models
- 用变体题目检验大模型与人类在类比推理中的抽象能力
- 大模型在字母串类比上性能大幅下降,而人类保持稳定
- 大模型对答案顺序和改写敏感,适合评估认知鲁棒性的研究
大语言模型在多个推理基准上表现良好,包括测试类比推理能力的任务。然而,其表现是真正的抽象推理,还是依赖于与预训练数据的相似性,仍存在争议。本文研究了三类由 Webb、Holyoak 与 Lu(2023)提出的类比任务的鲁棒性:字母串类比、数字矩阵类比和故事类比。对人类和 GPT 模型在原始问题及其变体上的表现进行对比,变体题保持相同抽象逻辑但与预训练数据差异较大。若系统具备稳健的抽象推理能力,性能不应显著下降。在简单字母串类比中,人类性能稳定,而 GPT 模型性能急剧下降;复杂度提升后,双方表现均差。数字矩阵类比中,仅一种变体导致模型性能下降。故事类比中,GPT 对答案顺序敏感,且更易受语义改写影响。结果表明,大模型在零样本类比推理中普遍缺乏鲁棒性,提示应同时关注准确率与鲁棒性来评估 AI 的认知能力。
原文摘要 · Abstract (English)
LLMs have performed well on several reasoning benchmarks, including ones that test analogical reasoning abilities. However, there is debate on the extent to which they are performing general abstract reasoning versus employing non-robust processes, e.g., that overly rely on similarity to pre-training data. Here we investigate the robustness of analogy-making abilities previously claimed for LLMs on three of four domains studied by Webb, Holyoak, and Lu (2023): letter-string analogies, digit matrices, and story analogies. For each domain we test humans and GPT models on robustness to variants of the original analogy problems that test the same abstract reasoning abilities but are likely dissimilar from tasks in the pre-training data. The performance of a system that uses robust abstract reasoning should not decline substantially on these variants. On simple letter-string analogies, we find that while the performance of humans remains high for two types of variants we tested, the GPT models' performance declines sharply. This pattern is less pronounced as the complexity of these problems is increased, as both humans and GPT models perform poorly on both the original and variant problems requiring more complex analogies. On digit-matrix problems, we find a similar pattern but only on one out of the two types of variants we tested. On story-based analogy problems, we find that, unlike humans, the performance of GPT models are susceptible to answer-order effects, and that GPT models also may be more sensitive than humans to paraphrasing. This work provides evidence that LLMs often lack the robustness of zero-shot human analogy-making, exhibiting brittleness on most of the variations we tested. More generally, this work points to the importance of carefully evaluating AI systems not only for accuracy but also robustness when testing their cognitive capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。