构建双语义变化数据集,提升俚语与标准用法的语义演变检测能力
The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

- 设计双向语义变化与俚语消歧双数据集,覆盖三时期语义轨迹
- GPT-4o在多标签准确率上表现最佳,但罕见俚语识别仍不足0.5
- 适合自然语言处理、语义演化研究者使用,尤其关注非正式语言
自动语义变化检测旨在识别词义随时间的演变,揭示语言与社会变迁。尽管计算词汇语义变化(LSC)取得进展,现有基准和方法难以捕捉双向语义变化,尤其是同时存在意义增减的情况,尤其在兼具俚语与标准用法的词语中更为突出。为弥补这一空白,我们构建两个互补的基准数据集:双向词汇语义变化(BD-LSC)数据集涵盖三个时期中的语义增益、损失与稳定,支持复杂语义轨迹分析;俚语追踪词义消歧(ST-WSD)数据集提供俚语与标准用法结合词的细粒度实例级语义标注,支持词义消歧与语义变化检测模型的系统评估。我们基于这些基准对不同方法体系进行系统评测:基于上下文嵌入的无监督聚类、有监督机器学习、基于Transformer的模型以及前沿大语言模型。结果表明,少样本GPT-4o在精确语义匹配(ESM)与多标签准确率上表现最优;然而所有系统在宏平均F1分数均接近0.5,显示稀有俚语语义仍属核心挑战。
原文摘要 · Abstract (English)
Automatic semantic change detection aims to identify how word meanings shift over time, offering insights into both linguistic and societal change. Despite recent progress in computational lexical semantic change (LSC), existing benchmarks and methods struggle to capture bi-directional semantic change, particularly cases where words simultaneously gain and lose senses. This problem is especially challenging for words that have both slang and standard meanings. To address these gaps, we introduce two complementary benchmark datasets. The Bi-Directional Lexical Semantic Change (BD-LSC) dataset captures sense gain, sense loss, and stability across three time periods, enabling the study of complex semantic trajectories. The SlangTrack Word Sense Disambiguation (ST-WSD) dataset provides fine-grained, instance-level sense annotations for words combining slang and standard usages, supporting systematic benchmarking of WSD and semantic change detection models. Using these benchmarks, we systematically evaluate models across different methodological families: unsupervised clustering using contextualised embeddings, supervised machine learning, transformer-based models, and state-of-the-art large language models. Among the evaluated systems, the few-shot GPT-4o model achieved the strongest aggregate performance on Exact Sense Match (ESM) and multi-label accuracy; however, Macro-F1 scores near 0.5 across all systems show that rare slang senses remain difficult, which we identify as the central open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。