arXiv:2410.17482cs.CL2024-10中稿 · GenBench 2024被引 6

大模型能理解新词组含义,但无法完全模仿人类判断分布。

Is artificial intelligence still intelligence? LLMs generalize to novel adjective-noun pairs, but don't mimic the full human distribution

  • 通过上下文推理,大模型可泛化到未见词组组合。
  • 在无上下文测试中,模型回答分布与人类相似度达75%。
  • 适合研究语言理解、认知模拟与模型评估的学者参考。

对“人工智能是否仍是智能”这类形容词-名词组合的推理,是检验大语言模型语义理解与组合泛化能力的良好测试场景,因为这些组合对人类和模型而言都是新颖的,却仍能引发一致的人类判断。我们测试了多种大语言模型,发现最大规模的模型在依赖上下文时能做出类人推理,并可泛化至未见的词组组合。我们还提出了三种在无上下文条件下评估模型推理的方法,此时人类答案呈分布状而非单一正确答案。结果显示,模型在最多75%的数据集上表现出类人分布,虽具潜力但仍存提升空间。

原文摘要 · Abstract (English)

Inferences from adjective-noun combinations like "Is artificial intelligence still intelligence?" provide a good test bed for LLMs' understanding of meaning and compositional generalization capability, since there are many combinations which are novel to both humans and LLMs but nevertheless elicit convergent human judgments. We study a range of LLMs and find that the largest models we tested are able to draw human-like inferences when the inference is determined by context and can generalize to unseen adjective-noun combinations. We also propose three methods to evaluate LLMs on these inferences out of context, where there is a distribution of human-like answers rather than a single correct answer. We find that LLMs show a human-like distribution on at most 75\% of our dataset, which is promising but still leaves room for improvement.

大模型语义理解泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。