arXiv:2510.13205cs.LGcs.AI2025-10被引 2

用专家规则指导模型检测医疗骗保行为,提升准确率与可解释性。

CleverCatch: A Knowledge-Guided Weak Supervision Model for Fraud Detection

  • 将医疗领域规则融入神经网络嵌入空间,实现规则与数据对齐。
  • 在真实数据集上比顶尖基线平均提升1.3% AUC和3.4%召回率。
  • 适合需高透明度的医疗风控场景,兼顾精准与可解释性。

医疗骗保检测因标注数据稀缺、欺诈手法不断演变及病历数据高维性而面临挑战。传统监督方法受限于极端标签稀疏,纯无监督方法则难以捕捉临床有意义的异常。本文提出CleverCatch,一种基于知识引导的弱监督模型,用于检测欺诈性处方行为,兼具更高准确率与可解释性。该方法将结构化领域知识嵌入神经架构,在共享嵌入空间中对齐规则与数据样本。通过联合训练合成的合规与违规数据编码器,模型学习到可泛化的软规则嵌入,使数据驱动学习受领域约束增强,弥合专家经验与机器学习之间的鸿沟。在大规模真实数据集上的实验表明,CleverCatch超越四种先进异常检测基线,平均提升1.3% AUC与3.4%召回率。消融研究进一步验证了专家规则的互补作用,证明该框架具备良好适应性。结果表明,将专家规则嵌入学习过程不仅能提升检测性能,还能增强透明度,为医疗骗保等高风险领域提供可解释的解决方案。

原文摘要 · Abstract (English)

Healthcare fraud detection remains a critical challenge due to limited availability of labeled data, constantly evolving fraud tactics, and the high dimensionality of medical records. Traditional supervised methods are challenged by extreme label scarcity, while purely unsupervised approaches often fail to capture clinically meaningful anomalies. In this work, we introduce CleverCatch, a knowledge-guided weak supervision model designed to detect fraudulent prescription behaviors with improved accuracy and interpretability. Our approach integrates structured domain expertise into a neural architecture that aligns rules and data samples within a shared embedding space. By training encoders jointly on synthetic data representing both compliance and violation, CleverCatch learns soft rule embeddings that generalize to complex, real-world datasets. This hybrid design enables data-driven learning to be enhanced by domain-informed constraints, bridging the gap between expert heuristics and machine learning. Experiments on the large-scale real-world dataset demonstrate that CleverCatch outperforms four state-of-the-art anomaly detection baselines, yielding average improvements of 1.3\% in AUC and 3.4\% in recall. Our ablation study further highlights the complementary role of expert rules, confirming the adaptability of the framework. The results suggest that embedding expert rules into the learning process not only improves detection accuracy but also increases transparency, offering an interpretable approach for high-stakes domains such as healthcare fraud detection.

欺诈检测弱监督医疗风控可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。