arXiv:2609.04804cs.AI2026-09

针对医疗时间序列生成中少数类特征被掩盖的问题,提出多尺度流匹配方法增强罕见病模式。

MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis

论文配图:MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis
图 1 · 摘自论文原文
  • 采用分层量化编码器在不同时间尺度上捕捉临床趋势与细微动态
  • 通过类别条件令牌引导提升少数类样本的生成质量,平均AUPRC提升5.8%
  • 适合处理不平衡医疗数据的合成,尤其对罕见病研究有实用价值

合成医疗时间序列可缓解数据稀缺问题,支持可靠临床预测模型的开发。然而现有方法主要关注整体分布和时序动态的匹配,未必能提升在不平衡医疗数据上的下游性能。临床关键模式常出现在异构时间尺度上,而少数类特征易被多数群体模式掩盖。为此,本文提出MedFlow,一种面向医疗时间序列合成的类别感知多尺度流匹配框架。MedFlow使用向量量化多尺度分词器,在互补的时间分辨率下表示医疗序列,同时捕捉宏观临床趋势与细粒度动态。进一步引入令牌边缘引导机制,将类别条件令牌统计信息直接融入流匹配过程,引导生成趋向特定类别的令牌空间区域,强化少数类模式的同时保持真实数据的整体与尾部分布。在四个公开数据集(涵盖电子健康记录、脑电图和心电图信号)上的实验表明,MedFlow在下游预测任务中持续优于近期最先进的基于扩散模型的基线方法。平均而言,其AUPRC提升5.8%,上下文FID降低88.6%,采样吞吐量提高3.8倍。

原文摘要 · Abstract (English)

Synthetic medical time-series generation can alleviate data scarcity and support the development of reliable clinical prediction models. However, existing methods mainly focus on matching the overall distribution and temporal dynamics of real data, which does not necessarily ensure strong downstream utility on imbalanced medical datasets. Clinically informative patterns often occur at heterogeneous temporal scales, while rare minority-class characteristics can be obscured by dominant population patterns. To address these challenges, we propose MedFlow, a class-aware multi-scale flow matching framework for medical time-series synthesis. MedFlow employs a vector-quantized multi-scale tokenizer to represent medical sequences at complementary temporal resolutions, capturing both coarse clinical trends and fine-grained dynamics. We further introduce Token Marginal Guidance, which incorporates class-conditional token statistics directly into the flow matching process to steer generation toward class-specific regions of the learned tokens. This mechanism strengthens minority-class patterns, while preserving the global and tail distributions of real data. Experiments on four public datasets covering electronic health records, EEG, and ECG signals demonstrate that MedFlow consistently outperforms recent state-of-the-art diffusion-based baselines across downstream prediction tasks. On average, it improves AUPRC by 5.8%, reduces Context-FID by 88.6%, and achieves 3.8$\times$ higher sampling throughput.

医疗生成时间序列多尺度少数类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。