跨30语言对比物体命名数据,提升可比性与透明度
Everybody Likes to Sleep: A Computer-Assisted Comparison of Object Naming Data from 30 Languages
- 用计算机辅助方法统一30语言的物体命名条目
- 发现多数语言中常见概念具有高度重叠性
- 适合语言学、认知科学及跨语言研究者参考
物体命名是人际交流中的基础能力,广泛应用于心理语言学、认知语言学及语言与视觉研究。现有命名数据集常缺乏透明度且结构各异。本研究通过多语言、计算机辅助方法,将30种语言、10个语系的17个命名数据集中的条目关联至统一概念。我们展示了如何通过搜索多数数据集中重复出现的概念,以及对比这些数据集与历史语言学和语言类型学中经典基本词汇表的概念空间。研究成果可为跨语言物体命名研究提供基础,并指导未来相关任务的设计。
原文摘要 · Abstract (English)
Object naming - the act of identifying an object with a word or a phrase - is a fundamental skill in interpersonal communication, relevant to many disciplines, such as psycholinguistics, cognitive linguistics, or language and vision research. Object naming datasets, which consist of concept lists with picture pairings, are used to gain insights into how humans access and select names for objects in their surroundings and to study the cognitive processes involved in converting visual stimuli into semantic concepts. Unfortunately, object naming datasets often lack transparency and have a highly idiosyncratic structure. Our study tries to make current object naming data transparent and comparable by using a multilingual, computer-assisted approach that links individual items of object naming lists to unified concepts. Our current sample links 17 object naming datasets that cover 30 languages from 10 different language families. We illustrate how the comparative dataset can be explored by searching for concepts that recur across the majority of datasets and comparing the conceptual spaces of covered object naming datasets with classical basic vocabulary lists from historical linguistics and linguistic typology. Our findings can serve as a basis for enhancing cross-linguistic object naming research and as a guideline for future studies dealing with object naming tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。