SuperMAN可直接处理异构稀疏时间数据,兼具高精度与可解释性。
SuperMAN: Interpretable and Expressive Networks over Temporally Sparse Heterogeneous Data
- 将异构时间信号建模为隐式图集合,实现端到端学习
- 在克罗恩病预测等任务中达顶尖性能,准确率显著提升
- 支持多层级可解释性,适合医疗等高风险场景
现实世界的时间数据通常包含多种信号类型,在不规则、异步的时间点上采样。例如,在医疗领域,不同血液检测项目在不同时间与频率下进行,导致数据碎片化且分布不均。类似问题也出现在大型系统监控的事件日志中。有效学习此类数据需处理异构、稀疏的时间信号。本文提出超混合加性网络(SuperMAN),一种设计即具可解释性的新框架,通过将信号建模为隐式图集合,直接从这类数据中学习。SuperMAN提供节点级、图级及子集级的重要性分析,支持在具备领域先验时权衡细粒度可解释性与更强表达能力。在真实世界高风险任务中表现卓越,包括基于常规血液检测预测克罗恩病发病和住院时长,以及虚假新闻检测。此外,其可解释性特性有助于揭示疾病发展阶段转变,为医疗领域提供关键洞察。
原文摘要 · Abstract (English)
Real-world temporal data often consists of multiple signal types recorded at irregular, asynchronous intervals. For instance, in the medical domain, different types of blood tests can be measured at different times and frequencies, resulting in fragmented and unevenly scattered temporal data. Similar issues of irregular sampling occur in other domains, such as the monitoring of large systems using event log files. Effectively learning from such data requires handling sets of temporal sparse and heterogeneous signals. In this work, we propose Super Mixing Additive Networks (SuperMAN), a novel and interpretable-by-design framework for learning directly from such heterogeneous signals, by modeling them as sets of implicit graphs. SuperMAN provides diverse interpretability capabilities, including node-level, graph-level, and subset-level importance, and enables practitioners to trade finer-grained interpretability for greater expressivity when domain priors are available. SuperMAN achieves state-of-the-art performance in real-world high-stakes tasks, including predicting Crohn's disease onset and hospital length of stay from routine blood test measurements and detecting fake news. Furthermore, we demonstrate how SuperMAN's interpretability properties assist in revealing disease development phase transitions and provide crucial insights in the healthcare domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。