arXiv:2505.21301cs.CLcs.AI2025-05ACL被引 4

对比人类与大模型在细粒度分类上的表现,发现差距明显但领域差异显著。

How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian

  • 用意大利语构建187个具体词的次级类别示例数据集
  • 大模型生成示例与人类组织方式对齐度低,跨语义域表现不一
  • 提示研究者慎用AI生成数据,尤其关注领域差异

人们可对同一实体在多个分类层级上进行归类,如基本层(熊)、上位层(动物)和次级层(灰熊)。以往研究多聚焦于基本层类别,本研究首次通过分析次级类别实例来探究概念知识的组织方式。我们构建了一个新的意大利语心理语言学数据集,包含187个具体词汇的由人类生成的次级类别示例。随后,利用这些数据评估文本与视觉大模型在三个关键任务中的表现:实例生成、类别归纳和典型性判断。结果表明,人类与大模型之间存在较低的一致性,与先前研究一致;但不同语义领域间的性能差异显著。该研究揭示了使用AI生成示例支持心理学与语言学研究的潜力与局限。

原文摘要 · Abstract (English)

People can categorize the same entity at multiple taxonomic levels, such as basic (bear), superordinate (animal), and subordinate (grizzly bear). While prior research has focused on basic-level categories, this study is the first attempt to examine the organization of categories by analyzing exemplars produced at the subordinate level. We present a new Italian psycholinguistic dataset of human-generated exemplars for 187 concrete words. We then use these data to evaluate whether textual and vision LLMs produce meaningful exemplars that align with human category organization across three key tasks: exemplar generation, category induction, and typicality judgment. Our findings show a low alignment between humans and LLMs, consistent with previous studies. However, their performance varies notably across different semantic domains. Ultimately, this study highlights both the promises and the constraints of using AI-generated exemplars to support psychological and linguistic research.

认知科学大模型评估语言习得

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。