arXiv:2602.22045cs.CL2026-02KDD

构建首个大规模区块链领域文本数据集,揭示科研领先于市场创新的规律

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

  • 整合2.98亿词元的科学文献、专利和社交媒体文本,覆盖超2200万条记录
  • 发现科技在学术论文中首次出现,随后才进入专利与社交讨论,形成正向循环
  • 适合区块链研究、金融科技分析及自然语言处理领域的学者与从业者

我们提出DLT-Corpus,迄今最大的分布式账本技术(DLT)专用文本集合:包含2.98亿词元,来自2212万篇文档,涵盖37,440篇科学文献、49,023件美国专利商标局(USPTO)专利以及2200万条社交媒体帖子。现有针对DLT的自然语言处理资源多集中于加密货币价格预测与智能合约,忽视了该领域语言特性的深度挖掘。尽管行业市值达约3万亿美元且技术快速演进,仍缺乏系统性语料支持。我们通过分析技术兴起模式与市场创新关联性,发现新技术通常先出现在科学文献中,再延伸至专利与社交网络,符合传统技术转移路径。尽管社交媒体情绪长期看涨,但科研与专利活动更关注长期趋势,与整体市场扩张形成良性循环。我们公开发布DLT-Corpus及配套资源:LedgerBERT(在特定命名实体识别任务上较BERT-base提升23%)、23,301条加密新闻标题与描述的情感分析数据集、工具与代码。

原文摘要 · Abstract (English)

We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents spanning scientific literature (37,440 publications), United States Patent and Trademark Office (USPTO) patents (49,023 filings), and social media (22 million posts). Existing Natural Language Processing (NLP) resources for DLT focus narrowly on cryptocurrency price prediction and smart contracts, leaving domain-specific language underexplored despite the sector's ~$3 trillion market capitalization and rapid technological evolution. We demonstrate DLT-Corpus' utility by analyzing patterns of technology emergence and market-innovation correlations. Findings reveal that technologies first appear in our scientific literature subset before reaching patents and social media, following traditional technology transfer patterns. While social media sentiment remains overwhelmingly bullish even during crypto winters, scientific and patent activity grows less tied to short-term sentiment, tracking overall market expansion in a virtuous cycle in which research precedes and enables economic growth that, in turn, funds further innovation. We release the DLT-Corpus and companion artifacts: LedgerBERT (+23% over BERT-base on DLT-specific Named Entity Recognition (NER) task), a sentiment analysis dataset of 23,301 crypto news headlines and descriptions, tools, and code.

区块链文本数据集NLP知识发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。