通过学习特征嵌入权重自动发现临床时间序列中的相关特征组。
Data-Driven Discovery of Feature Groups in Clinical Time Series
- 基于特征嵌入层权重聚类,动态发现与任务相关的特征分组。
- 在真实医疗数据上性能媲美专家定义的分组,且可解释性更强。
- 适合临床数据分析、深度学习模型可解释性研究者使用。
临床时间序列数据对患者监测和预测建模至关重要。这些数据通常是多变量的,包含来自不同来源的数百个异构特征。根据相似性和与预测任务的相关性对特征进行分组,已被证明能提升深度学习架构的性能。然而,仅依靠语义知识预先定义这些分组对领域专家也极具挑战。为此,我们提出一种新方法:通过聚类特征级嵌入层的权重来学习特征组。该方法可无缝集成到标准监督训练中,并发现直接提升下游临床任务性能的特征组。我们在合成数据上验证了该方法优于静态聚类,在真实世界医学数据上实现了与专家定义分组相当的性能。此外,学习到的特征组具有临床可解释性,支持数据驱动地发现变量间的任务相关关系。
原文摘要 · Abstract (English)
Clinical time series data are critical for patient monitoring and predictive modeling. These time series are typically multivariate and often comprise hundreds of heterogeneous features from different data sources. The grouping of features based on similarity and relevance to the prediction task has been shown to enhance the performance of deep learning architectures. However, defining these groups a priori using only semantic knowledge is challenging, even for domain experts. To address this, we propose a novel method that learns feature groups by clustering weights of feature-wise embedding layers. This approach seamlessly integrates into standard supervised training and discovers the groups that directly improve downstream performance on clinically relevant tasks. We demonstrate that our method outperforms static clustering approaches on synthetic data and achieves performance comparable to expert-defined groups on real-world medical data. Moreover, the learned feature groups are clinically interpretable, enabling data-driven discovery of task-relevant relationships between variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。