arXiv:2510.04622cs.LGeess.SP2025-10

用预测模型生成高保真生物医学时序数据,解决隐私与数据短缺难题。

Forecasting-based Biomedical Time-series Data Synthesis for Open Data and Robust AI

  • 基于最新预测模型构建合成数据框架,捕捉复杂电生理信号动态。
  • 睡眠阶段分类任务中合成数据使准确率提升至91.00%,增益达3.71%。
  • 适合需要开放数据共享与鲁棒AI训练的研究者使用。

由于严格的隐私法规和高昂的资源成本,生物医学时序数据获取受限,严重制约了该领域AI的发展,导致数据需求与可及性之间存在巨大缺口。合成数据生成通过创建保持真实生物医学时序数据统计特性的虚拟数据集,在不泄露患者隐私的前提下提供解决方案。尽管GAN、VAE和扩散模型能捕捉全局数据分布,但预测模型具备针对序列动态的归纳偏置。本文提出一种基于近期预测模型的生物医学时序数据合成框架,可高保真地还原如EEG和EMG等复杂电生理信号。这些合成数据可自由用于开放AI开发,并持续提升下游模型性能。在睡眠阶段分类任务中的数值结果显示,数据增强后性能最高提升3.71%,纯合成数据达到91.00%准确率,超过仅用真实数据的基线表现。

原文摘要 · Abstract (English)

The limited data availability due to strict privacy regulations and significant resource demands severely constrains biomedical time-series AI development, which creates a critical gap between data requirements and accessibility. Synthetic data generation presents a promising solution by producing artificial datasets that maintain the statistical properties of real biomedical time-series data without compromising patient confidentiality. While GANs, VAEs, and diffusion models capture global data distributions, forecasting models offer inductive biases tailored for sequential dynamics. We propose a framework for synthetic biomedical time-series data generation based on recent forecasting models that accurately replicates complex electrophysiological signals such as EEG and EMG with high fidelity. These synthetic datasets can be freely shared for open AI development and consistently improve downstream model performance. Numerical results on sleep-stage classification show up to a 3.71\% performance gain with augmentation and a 91.00\% synthetic-only accuracy that surpasses the real-data-only baseline.

时序生成合成数据生物医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。