用网页摘要提升中国供应链图谱构建效率,突破传统数据局限。
Snippet-Driven Supply Chain Discovery with LLMs: Scaling Visibility in China

- 用搜索摘要作初筛,降低大模型处理成本
- 覆盖7.2倍企业、9.3倍关系,远超传统披露数据
- 保留溯源信息,适合金融与经济研究者使用
金融与经济研究常依赖结构化的供应链披露和商业数据库。在中国,供应商-客户披露通常仅限上市公司主要合作方,导致非上市企业及长尾企业间联系在结构化数据中覆盖不足。公共网络证据可通过企业、政府及贸易媒体披露部分弥补此缺口,但大规模全文挖掘成本高昂,因页面常无法访问或需高成本处理。本文提出一种基于摘要的供应链知识图谱(SCKG)构建方法,以企业为节点、企业间关系为边。网页搜索摘要为搜索结果返回的查询偏倚摘要,作为大语言模型(LLM)关系抽取的可扩展第一层证据。评估表明:全文字块处理虽发现19.8倍更多唯一关系,但需251.2倍更多输入token且冗余更高;以130,685家中国公司为搜索种子(含2024年沪深上市及大型非上市企业),在上市公司子集上,所建SCKG覆盖企业达7.2倍、关系达9.3倍于基于CSMAR披露的基准,揭示重尾度分布特征。保留的溯源元数据使SCKG成为披露数据库的可审计补充。
原文摘要 · Abstract (English)
Financial and economic research often relies on structured supply-chain disclosures and commercial databases. In China, supplier--customer disclosure is typically limited to major partners of listed firms, leaving unlisted firms and long-tail inter-firm links poorly captured in structured data. Public web evidence can partly complement this gap through corporate, government, and trade-media disclosures; however, full-text web mining at scale is costly because pages are often inaccessible or expensive to process with large language models (LLMs). We propose a snippet-driven method for constructing a supply chain knowledge graph (SCKG), with firms as nodes and inter-firm relationships as edges. Web search snippets are query-biased summaries returned with search results. We use them as a scalable first-pass evidence layer for LLM-based relationship extraction. We evaluate the pipeline in terms of extraction efficiency and coverage. For extraction efficiency, exhaustive full-text chunking discovers 19.8$\times$ more unique relationships than snippets, but requires 251.2$\times$ more input tokens and yields higher redundancy. For coverage, we use 130,685 Chinese firms as search seeds, covering Shanghai/Shenzhen-listed firms and large unlisted firms as of 2024. In the listed-firm subset, the resulting SCKG covers 7.2$\times$ more firms and 9.3$\times$ more relationships than the CSMAR disclosure-based benchmark, while revealing heavy-tailed degree patterns. Retained provenance metadata make the SCKG an auditable complement to disclosure-based databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。