arXiv:2411.02348cs.AIcs.CL2024-11Transactions of th…被引 9

大模型解类比题像孩子一样迁移?实验发现它们做不到。

Can Large Language Models generalize analogy solving like children can?

  • 用字母串类比测试人类与大模型的跨域迁移能力
  • 儿童和成人能轻松迁移到希腊字母和符号域,大模型不行
  • 揭示大模型在类比推理上的泛化缺陷,适合研究认知差距者看

人类在童年时期便具备解决类比问题的能力,如“body : feet :: table : ?”,且能轻易迁移到其他领域,例如视觉类比“( : ) :: < : ?”。近期研究表明大语言模型(LLMs)可解决多种类比形式。但它们能否像人类一样在新领域泛化?我们让儿童、成人及大模型在拉丁字母、近迁移域(希腊字母)和远迁移域(符号列表)中解答一系列字母串类比(如 a b : a c :: j k : ?)。结果发现,儿童与成人能轻松将知识迁移到陌生领域,而大模型未能实现有效泛化。这一关键差异表明,当前大模型仍难以实现稳健的人类式类比迁移。

原文摘要 · Abstract (English)

In people, the ability to solve analogies such as "body : feet :: table : ?" emerges in childhood, and appears to transfer easily to other domains, such as the visual domain "( : ) :: < : ?". Recent research shows that large language models (LLMs) can solve various forms of analogies. However, can LLMs generalize analogy solving to new domains like people can? To investigate this, we had children, adults, and LLMs solve a series of letter-string analogies (e.g., a b : a c :: j k : ?) in the Latin alphabet, in a near transfer domain (Greek alphabet), and a far transfer domain (list of symbols). Children and adults easily generalized their knowledge to unfamiliar domains, whereas LLMs did not. This key difference between human and AI performance is evidence that these LLMs still struggle with robust human-like analogical transfer.

类比推理大模型认知研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。