arXiv:2605.17758cs.LG2026-05

Memisis用大模型自动生成并评估医疗表格数据,兼顾隐私与公平性。

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

论文配图:Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets
图 1 · 摘自论文原文
  • 通过大模型自然语言指令控制数据生成流程,自动化调用多种合成方法。
  • 在精神分裂症数据集上测试6种合成器,验证了生成数据在隐私、效用和公平性上的平衡。
  • 适合需要合规生成医疗数据的研究者或临床决策支持系统开发者。

合成数据在医疗领域广泛应用,可在不暴露敏感患者信息的前提下保留真实数据的统计特性。在隐私、效用和公平性维度上生成并评估合成数据,对下游预测任务和临床决策至关重要。我们提出Memisis,一个整合现有合成工具库、大语言模型(LLMs)和先进评估指标的工具,实现数据生成、验证与评估的统一工作流。用户可自定义训练规模、训练轮次及合成样本数量。除手动配置外,还支持交互式代理模式,用户以自然语言描述生成目标,系统自动调用合成器并完成评估。实验使用包含种族和性别等受保护属性的开源精神分裂症数据集,对比评估了涵盖GAN、VAE、扩散模型与归一化流的六种合成器,并利用本地LLM协调全流程。该系统赋予用户对生成与评估过程的高度灵活性与控制力。

原文摘要 · Abstract (English)

Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information. Generating and evaluating synthetic data across privacy, utility, and fairness dimensions is crucial for enabling high-quality data availability in downstream prediction tasks and clinical decision making. We present \textbf{Memisis}, a tool that orchestrates and evaluates synthetic data by leveraging existing synthesis libraries, large language models (LLMs), and state-of-the-art evaluation metrics. Our tool creates a unified workflow for data generation, validation, and evaluation. Users can control training size, training epochs, and the number of synthetic rows to sample. Beyond manual configuration, an interactive agent mode allows users to specify data generation goals in natural language, and the tool orchestrates the full pipeline by invoking existing synthesizers while performing the requisite evaluation. For the demo, we use an open-source schizophrenia dataset with protected attributes related to race and gender, evaluate six synthesizers spanning GANs, VAEs, diffusion models, and normalizing flows, and use a local LLM to orchestrate the workflow. The system affords users flexibility and control over the data generation and evaluation process.

合成数据医疗AI大模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。