分析德英语名词复合词随时间演变的语义透明度,发现其可预测性变化趋势。
Losing My Composure: Predicting Compositionality Over Time

- 构建跨年代语料库,标注23个德语与26个英语复合词的逐十年语义透明度。
- 发现复合词整体透明度仅呈轻微下降趋势,与学界普遍观点不符。
- 窄时窗训练模型比宽时窗更准确,静态表示也表现不俗,适合长期演化建模。
本文研究德语与英语名词复合词的语义演变现象,旨在探究和建模其意义及组合性随时间的渐进变化。为此,提出「组合性趋势预测」任务,并构建了一个新颖的历时语料库数据集,涵盖23个德语和26个英语目标复合词,提供跨越数十年的上下文组合性评分,实现逐十年度标注与趋势分析。该数据使我们能实证检验以往未被验证的假设,如复合词是否随时间变得越来越不透明。基于这些标注,我们对不同复杂度的语义向量表示进行了实验,采用多种时间粒度在历时数据上训练,共生成约100个每类表示模型,每个覆盖1至5个十年的时间段。结果表明,尽管文献中普遍认为组合性会显著下降,但本研究仅观察到微弱负向趋势。计算实验进一步显示,使用窄时间窗口(单十年或递增扩展窗口)训练的模型,相比在整个半个世纪跨度上训练的模型(即主流方法),更贴近实际的逐十年度评分。此外,静态表示在组合性趋势预测任务中表现与上下文表示相当。
原文摘要 · Abstract (English)
We explore the phenomenon of semantic change of German and English noun compounds, with the objective of investigating and modeling gradual changes of meanings and degrees of compositionality in the past and over time. To do so, we introduce the Compositionality Trend Prediction task, which is evaluated against a novel dataset of in-context compositionality ratings sampled across several decades of diachronic corpora for 23 German and 26 English target compounds, uniquely providing per-decade ratings and corresponding trends over time. These per-decade compositionality ratings allow us to investigate empirically untested hypotheses of generalized trends in compositionality over time, such as the idea that compounds should become less compositional (less transparent) over time. Beyond our empirical observations from the diachronic compositionality annotations, we perform experiments with semantic vector representations of varying complexity, as well as several temporal granularities for training these representations on diachronic data, resulting in about 100 models of each representation type, each covering a different 1--5 decade slice of a diachronic corpus. Contrary to the decisive tendency posited in the literature, we find only a small negative trend in compositionality over time in our target compounds. In our computational experiments, we find that using models trained on narrow time slices of diachronic data (single decades, or incrementally expanding temporal windows) align better with the per-decade compositionality ratings than those trained on an entire half-century window, the latter setting being an analog for the prevalent modeling approach of training representations on an entire half of a corpus' data. Additionally, we find static representations to be competitive with contextual representations in the Compositionality Trend Prediction task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。