用网页-知识-网页循环提升小企业供应商发现覆盖率。
Coverage-Aware Web Crawling for Domain-Specific Supplier Discovery via a Web--Knowledge--Web Pipeline
- 通过三阶段迭代:爬取网页、构建知识图谱、用图谱指导下一步爬取。
- 仅用144页就发现664家企业和542条关系,精度0.165,比基线少32%页面。
- 引入生态学方法评估覆盖度,适合供应链研究与工业数据分析者。
在特定行业领域中识别中小企业(SME)的完整图景对供应链韧性至关重要,但现有商业数据库存在显著覆盖缺口,尤其针对次级供应商和新兴细分市场企业。本文提出一种网页-知识-网页(W→K→W)迭代管道:首先爬取领域相关网页发现候选供应商实体;其次利用领域适配的少样本大模型提示技术,提取并整合结构化知识,构建异构知识图谱;最后利用知识图谱的拓扑结构与覆盖信号,引导后续爬取以填补供应商空间中的薄弱区域。为量化发现完整性,我们引入受生态学物种丰富度估计器(Chao1, ACE)启发的覆盖度评估框架。在半导体设备制造领域(NAICS 333242)的实验表明,该方法在仅使用144页的情况下,达到最高精度0.165和F1值0.123,较213页的基线减少32%的页面访问量,构建了包含664个实体和542条关系的知识图谱,且关系类型一致性达100%。
原文摘要 · Abstract (English)
Identifying the full landscape of small and medium-sized enterprises (SMEs) in specialized industry sectors is critical for supply-chain resilience, yet existing business databases suffer from substantial coverage gaps -- particularly for sub-tier suppliers and firms in emerging niche markets. We propose a \textbf{Web--Knowledge--Web (W$\to$K$\to$W)} pipeline that iteratively (1)~crawls domain-specific web sources to discover candidate supplier entities, (2)~extracts and consolidates structured knowledge into a heterogeneous knowledge graph using domain-adapted few-shot LLM prompting, and (3)~uses the knowledge graph's topology and coverage signals to guide subsequent crawling toward under-represented regions of the supplier space. To quantify discovery completeness, we introduce a \textbf{coverage estimation framework} inspired by ecological species-richness estimators (Chao1, ACE) adapted for web-entity populations. Experiments on the semiconductor equipment manufacturing sector (NAICS 333242) demonstrate that the W$\to$K$\to$W pipeline achieves the highest precision (0.165) and F1 (0.123) among all methods while using only 144 pages -- 32\% fewer than the 213-page baseline budget -- building a knowledge graph of 664 entities and 542 relations with 100\% relation type-consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。