arXiv:2508.11828cs.CL2025-08综述被引 3

梳理53个习语数据集,揭示心理语言学与计算研究的差异与空白

A Survey of Idiom Datasets for Psycholinguistic and Computational Research

  • 整合心理语言学与计算语言学中的习语数据集,分析标注方式与用途
  • 发现53个数据集中语言覆盖和任务类型扩展,但两大领域仍脱节
  • 适合从事语言认知、自然语言处理或跨学科研究的研究者参考

习语是意义无法从字面推断的隐喻表达,对计算处理和人类实验研究均构成挑战。本文综述了心理语言学与计算语言学领域为研究习语而开发的53个数据集,聚焦其内容、形式与使用目的。心理语言学资源通常包含熟悉度、透明度和构词性等维度的规范评分,而计算数据集则支持习语识别、分类、改写及跨语言建模等任务。本文总结了标注实践、覆盖范围与任务设计的趋势。尽管近年研究扩展了语言覆盖与任务多样性,但心理语言学与计算研究之间尚未建立关联。

原文摘要 · Abstract (English)

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets developed in psycholinguistics and computational linguistics for studying idioms, focusing on their content, form, and intended use. Psycholinguistic resources typically contain normed ratings along dimensions such as familiarity, transparency, and compositionality, while computational datasets support tasks like idiomaticity detection/classification, paraphrasing, and cross-lingual modeling. We present trends in annotation practices, coverage, and task framing across 53 datasets. Although recent efforts expanded language coverage and task diversity, there seems to be no relation yet between psycholinguistic and computational research on idioms.

习语研究数据集综述心理语言学计算语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。