arXiv:2608.19801cs.LG2026-08

用流匹配检测金融异常数据,关键在选对异常评分方法。

Unsupervised Anomaly Detection Using Flow Matching on Tabular Data

  • 用流匹配建模表格数据分布,避免依赖干净训练集
  • 轨迹型异常评分比单步评分更稳定,抗污染能力强
  • 森林流模型表现媲美甚至超过传统流匹配方法

金融异常检测常依赖大量未标注的交易日志,其中异常样本可能已在训练集中出现,违反了多数异常检测方法所依赖的“正常数据纯净”假设。尽管流匹配在生成建模中表现优异,但其在无监督表格数据异常检测中的鲁棒性仍待探索。本文通过对比时间条件收缩匹配(TCCM)与森林流(Forest-Flow),评估多种异常评分函数,发现评分选择至关重要:TCCM原用的单步决策评分对数据污染敏感,而基于轨迹的偏离度与重构评分则提供更稳定的异常信号。使用这些评分后,森林流在多个数据集上达到甚至超越TCCM性能。结果表明,在严重类别不平衡的金融场景下,异常评分机制对流匹配方法的成功尤为关键。

原文摘要 · Abstract (English)

Financial anomaly detection often relies on large unlabeled transaction logs, where anomalous samples may already be present during training. Such training-set contamination violates the clean-normal data assumption underlying many anomaly detection methods. Although flow matching has demonstrated strong performance in generative modeling, its robustness in unsupervised tabular anomaly detection remains underexplored. In this work, we study flow-matching-based anomaly detection under contaminated training data by comparing Time-Conditioned Contraction Matching (TCCM) with Forest-Flow and evaluating multiple anomaly scoring functions. Our results show that the choice of anomaly score is critical. The original single-step Decision score used by TCCM is sensitive to contamination, whereas trajectory-based Deviation and Reconstruction scores provide more stable anomaly signals. With these scores, Forest-Flow becomes competitive with, and in some cases outperforms, TCCM. These findings highlight the importance of anomaly scoring for flow-matching methods in financial anomaly detection under severe class imbalance.

异常检测流匹配表格数据金融风控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。