整合时间维度文本分析工具,让动态语义研究更便捷。
ttta: Tools for Temporal Text Analysis
- 构建统一工具包,支持跨时间的文本分析
- 解决传统方法忽略语言随时间演变的问题
- 适合研究社会、政治、经济等时序文本的学者
文本数据具有天然的时间属性,词语和短语的含义随时间变化,使用语境持续演化。这不仅体现在社交媒体中受热点、迷因和趋势快速影响的语言,也存在于新闻、经济或政治文本中。然而,大多数自然语言处理技术将语料库视为时间均质的整体,这种简化可能导致结果偏差。例如,在涵盖数年的语料上运行经典潜在狄利克雷分配(LDA),只能呈现整个时间段的“平均”主题分布,无法捕捉主题的动态演变。尽管已有多种时间文本分析工具,但它们分散在不同包和库中,难以统一、可复现地使用。ttta 包旨在作为一套集成工具,用于高效分析随时间演化的文本数据。
原文摘要 · Abstract (English)
Text data is inherently temporal. The meaning of words and phrases changes over time, and the context in which they are used is constantly evolving. This is not just true for social media data, where the language used is rapidly influenced by current events, memes and trends, but also for journalistic, economic or political text data. Most NLP techniques however consider the corpus at hand to be homogenous in regard to time. This is a simplification that can lead to biased results, as the meaning of words and phrases can change over time. For instance, running a classic Latent Dirichlet Allocation on a corpus that spans several years is not enough to capture changes in the topics over time, but only portraits an "average" topic distribution over the whole time span. Researchers have developed a number of tools for analyzing text data over time. However, these tools are often scattered across different packages and libraries, making it difficult for researchers to use them in a consistent and reproducible way. The ttta package is supposed to serve as a collection of tools for analyzing text data over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。