arXiv:2504.14808cs.CLcs.AI2025-04中稿 · the 2025 25th Inte…被引 1

通过邻近词信息动态优化词向量,提升特定领域文本表示效果。

On Self-improving Token Embeddings

  • 基于上下文邻近词迭代更新词向量,无需依赖大模型。
  • 在气象灾害文本中显著改善术语表征,捕捉灾情演变特征。
  • 适合领域文本分析、概念搜索等场景,尤其适用于小众词汇。

本文提出一种新颖且快速的静态词向量(或更一般地,标记向量)精炼方法。通过引入文本语料中邻近标记的嵌入信息,持续更新每个标记的表示,包括那些未预设嵌入的标记。该方法有效解决了词汇外问题。其运行独立于大型语言模型和浅层神经网络,可广泛应用于语料探索、概念搜索和词义消歧。该方法专为话题同质性较强的语料设计,在特定领域词汇受限的情况下,生成比通用预训练向量更有意义的嵌入。以美国国家海洋与大气管理局(NOAA)风暴事件数据库的子集为例,展示了该方法在探索风暴事件及其对基础设施与社区影响方面的应用。结果表明,该方法能随时间改进与风暴相关术语的表示,揭示灾难叙事的演变特征。

原文摘要 · Abstract (English)

This article introduces a novel and fast method for refining pre-trained static word or, more generally, token embeddings. By incorporating the embeddings of neighboring tokens in text corpora, it continuously updates the representation of each token, including those without pre-assigned embeddings. This approach effectively addresses the out-of-vocabulary problem, too. Operating independently of large language models and shallow neural networks, it enables versatile applications such as corpus exploration, conceptual search, and word sense disambiguation. The method is designed to enhance token representations within topically homogeneous corpora, where the vocabulary is restricted to a specific domain, resulting in more meaningful embeddings compared to general-purpose pre-trained vectors. As an example, the methodology is applied to explore storm events and their impacts on infrastructure and communities using narratives from a subset of the NOAA Storm Events database. The article also demonstrates how the approach improves the representation of storm-related terms over time, providing valuable insights into the evolving nature of disaster narratives.

词向量优化领域适配动态嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。