用熵衡量企业知识分布,发现关键实体难检索。
Business Entity Entropy
- 提出实体熵概念,量化知识在文档中的分散程度
- 发现熵呈重尾分布,实体越大熵越高
- 适合改进企业知识检索系统的设计
组织在多个平台生成海量互联内容。尽管语言模型可支持复杂推理,但从组织记忆中检索和上下文化信息仍具挑战。我们从熵的角度探索这一问题,提出实体熵度量,用于量化实体知识在文档中的分布,并设计一种受扩散模型启发的生成模型,以解释观察到的行为。在大规模企业语料库上的实证分析显示,熵分布呈重尾特征,实体规模与熵正相关,且不同类别存在特定熵模式。这些发现表明,并非所有实体都同样可检索,因而需要针对部分而非全部实体采用以实体为中心的检索或预处理策略。我们讨论了其实际意义及理论模型,以指导更高效的知识检索系统设计。
原文摘要 · Abstract (English)
Organizations generate vast amounts of interconnected content across various platforms. While language models enable sophisticated reasoning for use in business applications, retrieving and contextualizing information from organizational memory remains challenging. We explore this challenge through the lens of entropy, proposing a measure of entity entropy to quantify the distribution of an entity's knowledge across documents as well as a novel generative model inspired by diffusion models in order to provide an explanation for observed behaviours. Empirical analysis on a large-scale enterprise corpus reveals heavy-tailed entropy distributions, a correlation between entity size and entropy, and category-specific entropy patterns. These findings suggest that not all entities are equally retrievable, motivating the need for entity-centric retrieval or pre-processing strategies for a subset of, but not all, entities. We discuss practical implications and theoretical models to guide the design of more efficient knowledge retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。