arXiv:2502.10453cs.CRcs.AI2025-02中稿 · Financial Cryptogr…被引 3

用大模型将加密资产标签关联到知识图谱实体,提升溯源准确性。

Linking Cryptoasset Attribution Tags to Knowledge Graph Entities: An LLM-based Approach

  • 基于大模型构建端到端管道,自动匹配标签与知识图谱实体。
  • 在三个数据集上F1最高提升37.4%,召回率达93%且无需标注数据。
  • 本地模型性能接近远程模型,成本可降低90%仅损失1%效果。

Attribution tags 是现代加密资产溯源的基础,但标签不一致或错误可能导致调查误导甚至误判。为此,我们提出一种基于大语言模型(LLMs)的新计算方法,用于将标签与定义明确的知识图谱概念关联。我们实现了一个端到端流程,在三个公开可用的标签数据集上实验显示,该方法在F1得分上相比基线最高提升37.4%。通过引入概念过滤与阻断机制,生成包含五个知识图谱实体的候选集,实现93%的召回率,且无需标注数据。此外,我们验证了本地LLM模型在F1得分上可达90%,接近远程模型的94%表现。我们还分析了不同LLM与提示模板的成本-性能权衡,发现选择最优配置可降低90%成本,仅导致1%性能下降。本方法不仅提升了标签质量,也为构建更可靠的司法取证证据提供了范式。

原文摘要 · Abstract (English)

Attribution tags form the foundation of modern cryptoasset forensics. However, inconsistent or incorrect tags can mislead investigations and even result in false accusations. To address this issue, we propose a novel computational method based on Large Language Models (LLMs) to link attribution tags with well-defined knowledge graph concepts. We implemented this method in an end-to-end pipeline and conducted experiments showing that our approach outperforms baseline methods by up to 37.4% in F1-score across three publicly available attribution tag datasets. By integrating concept filtering and blocking procedures, we generate candidate sets containing five knowledge graph entities, achieving a recall of 93% without the need for labeled data. Additionally, we demonstrate that local LLM models can achieve F1-scores of 90%, comparable to remote models which achieve 94%. We also analyze the cost-performance trade-offs of various LLMs and prompt templates, showing that selecting the most cost-effective configuration can reduce costs by 90%, with only a 1% decrease in performance. Our method not only enhances attribution tag quality but also serves as a blueprint for fostering more reliable forensic evidence.

区块链溯源大模型应用知识图谱加密资产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。