arXiv:2505.01311cs.CL2025-05中稿 · the 2025 Annual Me…被引 1

用概率模型解析模糊时间副词在不同事件中的语义差异。

A Factorized Probabilistic Model of the Semantics of Vague Temporal Adverbials Relative to Different Event Types

  • 将模糊时间副词建模为可分解的概率分布,结合事件特性生成上下文语义。
  • 相比单一高斯模型,新模型在预测效果相近的前提下更简洁且易扩展。
  • 适合对自然语言时间理解、语义建模感兴趣的学者与开发者。

模糊时间副词(如recently、just、a long time ago)描述过去事件与说话时间之间的时距,但不明确具体时长。本文提出一种可分解的概率模型,将这些副词的语义表示为概率分布,并与特定事件的概率分布结合,生成上下文相关的语义解释。通过拟合母语者对不同事件发生时间与副词适用性判断的数据,我们发现该模型在预测能力上与基于单个高斯分布的非分解模型相当,但在奥卡姆剃刀原则下更具优势:结构更简单,且易于推广至新事件类型。

原文摘要 · Abstract (English)

Vague temporal adverbials, such as recently, just, and a long time ago, describe the temporal distance between a past event and the utterance time but leave the exact duration underspecified. In this paper, we introduce a factorized model that captures the semantics of these adverbials as probabilistic distributions. These distributions are composed with event-specific distributions to yield a contextualized meaning for an adverbial applied to a specific event. We fit the model's parameters using existing data capturing judgments of native speakers regarding the applicability of these vague temporal adverbials to events that took place a given time ago. Comparing our approach to a non-factorized model based on a single Gaussian distribution for each pair of event and temporal adverbial, we find that while both models have similar predictive power, our model is preferable in terms of Occam's razor, as it is simpler and has better extendability.

时间语义概率模型自然语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。