用词典与频率结合法压缩形式上下文,加速概念层次提取。
Reducing Formal Context Extraction: A Newly Proposed Framework from Big Corpora
- 结合WordNet和词频统计,减少形式上下文规模。
- 压缩后概念格保留98%原始层次质量,结构一致。
- 适合需快速构建概念体系的工业文本分析场景。
从自由文本自动抽取概念层次具有重要价值,但传统流程涉及多阶段处理,生成的形式上下文常含冗余或错误配对,导致计算耗时。为降低模糊性并提升效率,本文提出一种新框架,通过融合WordNet与词频方法压缩形式上下文规模。基于维基百科385个样本测试显示,该方法能有效缩减上下文大小,同时利用概念格不变量对比,结果概念格与标准格之间的同构性保持高达98%的层次质量。结构分析表明,压缩后的概念格保留了标准格的连接关系。此外,在不同密度的随机数据集上与五种基线方法对比,本方法在概念格性能上始终优于其他策略,且运行时间更优。
原文摘要 · Abstract (English)
Automating the extraction of concept hierarchies from free text is advantageous because manual generation is frequently labor- and resource-intensive. Free result, the whole procedure for concept hierarchy learning from free text entails several phases, including sentence-level text processing, sentence splitting, and tokenization. Lemmatization is after formal context analysis (FCA) to derive the pairings. Nevertheless, there could be a few uninteresting and incorrect pairings in the formal context. It may take a while to generate formal context; thus, size reduction formal context is necessary to weed out irrelevant and incorrect pairings to extract the concept lattice and hierarchies more quickly. This study aims to propose a framework for reducing formal context in extracting concept hierarchies from free text to reduce the ambiguity of the formal context. We achieve this by reducing the size of the formal context using a hybrid of a WordNet-based method and a frequency-based technique. Using 385 samples from the Wikipedia corpus and the suggested framework, tests are carried out to examine the reduced size of formal context, leading to concept lattice and concept hierarchy. With the help of concept lattice-invariants, the generated formal context lattice is compared to the normal one. In contrast to basic ones, the homomorphic between the resultant lattices retains up to 98% of the quality of the generating concept hierarchies, and the reduced concept lattice receives the structural connection of the standard one. Additionally, the new framework is compared to five baseline techniques to calculate the running time on random datasets with various densities. The findings demonstrate that, in various fill ratios, hybrid approaches of the proposed method outperform other indicated competing strategies in concept lattice performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。