arXiv:2604.23494cs.AIcs.LG2026-04

对比交易与地址级反洗钱评分差异,发现不同粒度导致审查结果大不相同。

Do Transaction-Level and Actor-Level AML Queues Agree? An Empirical Evaluation of Granularity Effects on the Elliptic++ Graph

  • 用四种聚合方式将交易级评分映射到地址级,构建可比审查队列。
  • 在1%审查预算下,动态评估时队列重合度仅0.374,静态评估更低至0.087。
  • 地址级评分更具时效性,少数时段中每百次审查识别出超91%非法行为。

基于图的区块链反洗钱系统可在交易或地址两个粒度上评分,但合规操作以地址为单位执行。本文提出一种评估方法,分析评分粒度如何影响固定审查预算下的调查队列组成。通过四类聚合算子将交易级评分投影至地址级动作单元,并引入预算化评估指标:收益率@预算、负担分解与案例碎片化。使用公开的Elliptic++比特币数据集(203,769笔交易;822,942个地址出现次数),在因果时间协议下分别训练独立随机森林分类器,通过杰卡德重合度、负担分解与特征匹配消融实验比较审查队列。在1%预算下,时序评估的均值杰卡德为0.374(标准差0.171);静态合并评估为0.087(95%置信区间[0.079, 0.094])。接收全部237个特征的地址模型重合度更低(杰卡德=0.051),每百次审查中非法占比仅4.3%,而交易投影队列为30.2%。地址级检测价值具有时序集中性:仅两个时间步超过91%非法率,而静态负担仅为3.4%。固定混合策略表现低于最优单粒度队列5.05个百分点(置信区间[-10.2pp, -0.9pp])。研究证实,评分粒度是反洗钱调查系统的关键设计变量——相同数据、相同预算,不同粒度导致不同审查对象。

原文摘要 · Abstract (English)

Graph-based anti-money laundering (AML) systems on blockchain networks can score suspicious activity at two granularity levels -- transactions or actor addresses -- yet compliance action is conducted per actor. This paper contributes an evaluation methodology for measuring how scoring granularity affects investigation queue composition under fixed review budgets. We formalize the evaluation through a projection framework mapping transaction-level scores to the actor-level action unit via four aggregation operators, and introduce budgeted investigation metrics -- yield@budget, burden decomposition, and case fragmentation. Using the public Elliptic++ Bitcoin dataset (203,769 transactions; 822,942 address occurrences), we train independent random forest classifiers at each level under a causal temporal protocol and compare review queues through Jaccard overlap, burden decomposition, and feature-matching ablations. At one-percent budget, temporal evaluation yields mean Jaccard of 0.374 (SD 0.171); static pooled evaluation yields 0.087 (95% CI [0.079, 0.094]). An enriched address model receiving all 237 features produces even lower overlap (Jaccard=0.051), with 4.3% illicit per 100 reviews versus 30.2% for the transaction-projected queue. Address-level detection value is temporally concentrated: two timesteps exceed 91% illicit per 100 reviews while the static burden is only 3.4%. A fixed hybrid policy underperforms the best single-level queue by 5.05pp (CI [-10.2pp, -0.9pp]). These findings establish that scoring granularity is a consequential design variable for AML investigation systems -- same data, same budget, different queues, different addresses investigated.

反洗钱图神经网络区块链评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。