构建通用实体关联挖掘框架,助力金融等领域文本分析与知识发现。
Building Entity Association Mining Framework for Knowledge Discovery
- 设计可插拔的实体提取与关联挖掘流水线,支持多源技术融合。
- 通过共现图与频率统计,量化识别实体间潜在关系与业务热点。
- 已在品牌产品发现和供应商风险监控中验证,适合快速构建文本应用。
从非结构化文本中提取有用信号以支持重要商业决策(如分析投资产品热度、客户偏好、风险监控等)是一项挑战。捕捉实体或概念间的交互关系并进行关联挖掘,是文本挖掘中的关键环节,有助于信息抽取、推理及知识发现,并可用于丰富或过滤知识图谱,引导探索过程、描述性分析,揭示文本中的隐藏规律。本文提出一种领域无关的通用框架,作为文档筛选、实体提取(支持DBpedia Spotlight、Spacy NER、自定义匹配器、词组/字典提取等插件)及关联关系挖掘的统一管道,可支撑各类文本挖掘业务场景,并提供量化评分指标用于排序。该框架包含三个核心组件:(a) 文档筛选:从海量文本中提取目标内容;(b) 可配置实体提取流水线;(c) 关联关系挖掘:生成共现图并基于共现频次统计,提供全局视角观察特定业务情境下的关联趋势或关注度。论文在品牌产品发现与供应商风险监控两个金融场景中展示了该框架的应用价值。我们期望该框架能减少重复工作,降低开发成本,推动关联挖掘应用的可复用性与快速原型设计。
原文摘要 · Abstract (English)
Extracting useful signals or pattern to support important business decisions for example analyzing investment product traction and discovering customer preference, risk monitoring etc. from unstructured text is a challenging task. Capturing interaction of entities or concepts and association mining is a crucial component in text mining, enabling information extraction and reasoning over and knowledge discovery from text. Furthermore, it can be used to enrich or filter knowledge graphs to guide exploration processes, descriptive analytics and uncover hidden stories in the text. In this paper, we introduce a domain independent pipeline i.e., generalized framework to enable document filtering, entity extraction using various sources (or techniques) as plug-ins and association mining to build any text mining business use-case and quantitatively define a scoring metric for ranking purpose. The proposed framework has three major components a) Document filtering: filtering documents/text of interest from massive amount of texts b) Configurable entity extraction pipeline: include entity extraction techniques i.e., i) DBpedia Spotlight, ii) Spacy NER, iii) Custom Entity Matcher, iv) Phrase extraction (or dictionary) based c) Association Relationship Mining: To generates co-occurrence graph to analyse potential relationships among entities, concepts. Further, co-occurrence count based frequency statistics provide a holistic window to observe association trends or buzz rate in specific business context. The paper demonstrates the usage of framework as fundamental building box in two financial use-cases namely brand product discovery and vendor risk monitoring. We aim that such framework will remove duplicated effort, minimize the development effort, and encourage reusability and rapid prototyping in association mining business applications for institutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。