通过消息处理时长构建特征,实现支付系统故障的早期定位与业务影响分析。
A Feature Engineering Approach for Business Impact-Oriented Failure Detection in Distributed Instant Payment Systems
- 基于连续ISO 20022消息间处理时长生成系统状态特征
- 在TIPS系统上成功检测多种异常模式并缩短排查时间
- 可解释性强,帮助运维人员理解故障的业务影响
即时支付基础设施每日需处理数百万笔交易,且要求零中断。传统监控手段难以将技术指标与业务可见性衔接。本文提出一种新型特征工程方法,基于连续ISO 20022消息交换的处理时长,构建系统状态紧凑表示。通过对此类特征应用异常检测,实现故障的早期发现与定位,并支持事件分类。在TARGET即时清算系统(TIPS)上的实验评估,结合真实故障与受控仿真,验证了该方法对多种异常模式的有效性,并提供内在可解释性,使运维人员能理解故障的业务影响。通过将特征映射至不同处理阶段,该框架可区分系统内外部问题,显著降低调查时间,弥合分布式系统中交易状态跨实体碎片化带来的可观测性缺口。
原文摘要 · Abstract (English)
Instant payment infrastructures have stringent performance requirements, processing millions of transactions daily with zero-downtime expectations. Traditional monitoring approaches fail to bridge the gap between technical infrastructure metrics and business process visibility. We introduce a novel feature engineering approach based on processing times computed between consecutive ISO 20022 message exchanges, creating a compact representation of system state. By applying anomaly detection to these features, we enable early failure detection and localization, allowing incident classification. Experimental evaluation on the TARGET Instant Payment Settlement (TIPS) system, using both real-world incidents and controlled simulations, demonstrates the approach's effectiveness in detecting diverse anomaly patterns and provides inherently interpretable explanations that enable operators to understand the business impact. By mapping features to distinct processing phases, the resulting framework differentiates between internal and external payment system issues, significantly reduces investigation time, and bridges observability gaps in distributed systems where transaction state is fragmented across multiple entities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。