研究语义边对文本网络统计特性的双重影响,帮人选对分析指标。
Probing the statistical properties of enriched co-occurrence networks
- 用词嵌入生成虚拟边,增强短文本的共现网络
- 虚拟边提升路径长度与中心性指标,但降低聚类系数有效性
- 停用词会改变网络性质,提示按文本长度选指标
近期研究尝试通过词嵌入为词共现网络添加虚拟边,以增强短文本的图表示。尽管这类增强网络已见成效,但语义边对传统共现网络的影响仍不明确。本文探究两类文本网络模型的统计特性:一是网络指标能否区分无意义与有意义文本;二是这些指标更敏感于句法还是语义层面。结果表明,引入虚拟边的效果因指标而异:在短文本中,平均最短路径和紧密中心性的信息量提升,但聚类系数的信息量随虚拟边增多而下降。此外,停用词的存在会影响增强网络的统计特性。研究结果可为不同文本长度和任务场景下选择合适的网络指标提供参考。
原文摘要 · Abstract (English)
Recent studies have explored the addition of virtual edges to word co-occurrence networks using word embeddings to enhance graph representations, particularly for short texts. While these enriched networks have demonstrated some success, the impact of incorporating semantic edges into traditional co-occurrence networks remains uncertain. This study investigates two key statistical properties of text-based network models. First, we assess whether network metrics can effectively distinguish between meaningless and meaningful texts. Second, we analyze whether these metrics are more sensitive to syntactic or semantic aspects of the text. Our results show that incorporating virtual edges can have positive and negative effects, depending on the specific network metric. For instance, the informativeness of the average shortest path and closeness centrality improves in short texts, while the clustering coefficient's informativeness decreases as more virtual edges are added. Additionally, we found that including stopwords affects the statistical properties of enriched networks. Our results can serve as a guideline for determining which network metrics are most appropriate for specific applications, depending on the typical text size and the nature of the problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。