arXiv:2503.11377cs.CLcs.DB2025-03被引 1

升级跨语言同义词数据库,提升数据质量与覆盖范围。

Advancing the Database of Cross-Linguistic Colexifications with New Workflows and Data

  • 构建新版本共指词数据库,全量标注语音转写。
  • 覆盖更多语系,样本更均衡,数据质量显著提升。
  • 适合语言类型学、心理语言学等领域的研究者使用。

词汇资源对跨语言分析至关重要,可为自然语言学习的计算模型提供新视角。本文介绍一个用于多义词比较研究(即共指现象)的先进数据库新版本。新版本在数据处理、筛选与展示方面均有改进。相较于以往版本,新数据库提供了更均衡的全球语言家族覆盖,数据质量更高,所有词形均以音标转写呈现。我们得出结论:新的跨语言共指词数据库有望激发关于语言类型学、历史语言学、心理语言学和计算语言学中开放问题的新研究。

原文摘要 · Abstract (English)

Lexical resources are crucial for cross-linguistic analysis and can provide new insights into computational models for natural language learning. Here, we present an advanced database for comparative studies of words with multiple meanings, a phenomenon known as colexification. The new version includes improvements in the handling, selection and presentation of the data. We compare the new database with previous versions and find that our improvements provide a more balanced sample covering more language families worldwide, with enhanced data quality, given that all word forms are provided in phonetic transcription. We conclude that the new Database of Cross-Linguistic Colexifications has the potential to inspire exciting new studies that link cross-linguistic data to open questions in linguistic typology, historical linguistics, psycholinguistics, and computational linguistics.

语言学共指词数据库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。