arXiv:2505.10718cs.CLcs.AI2025-05被引 4

用AI增强人类语义特征数据,提升覆盖率和预测精度。

AI-enhanced semantic feature norms for 786 concepts

  • 结合人类与大模型生成语义特征,提升数据密度
  • 在语义相似性预测上优于人工数据集和词嵌入模型
  • 验证后证明大模型可有效辅助认知科学研究

语义特征规范是研究人类概念知识的基础,但传统方法因依赖人力而面临概念/特征覆盖范围与质量可验证性之间的权衡。本文提出一种新方法,将大语言模型(LLMs)生成的响应与人类生成的特征规范数据结合,并通过可靠的人类判断验证其质量。结果表明,所构建的AI增强型语义特征数据集NOVA:Norms Optimized Via AI具有更高的特征密度和概念间重叠度,且在预测人类语义相似性判断方面优于同类人工数据集及词嵌入模型。研究显示,人类概念知识比以往数据集捕捉得更丰富,且经适当验证后,大模型可成为认知科学研究的强大工具。

原文摘要 · Abstract (English)

Semantic feature norms have been foundational in the study of human conceptual knowledge, yet traditional methods face trade-offs between concept/feature coverage and verifiability of quality due to the labor-intensive nature of norming studies. Here, we introduce a novel approach that augments a dataset of human-generated feature norms with responses from large language models (LLMs) while verifying the quality of norms against reliable human judgments. We find that our AI-enhanced feature norm dataset, NOVA: Norms Optimized Via AI, shows much higher feature density and overlap among concepts while outperforming a comparable human-only norm dataset and word-embedding models in predicting people's semantic similarity judgments. Taken together, we demonstrate that human conceptual knowledge is richer than captured in previous norm datasets and show that, with proper validation, LLMs can serve as powerful tools for cognitive science research.

语义特征大模型认知科学数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。