arXiv:2508.02771cs.LG2025-08中稿 · CIBB 2025 as a sho…

用合成数据解决医疗隐私问题,提升创伤机制分类效果

Synthetic medical data generation: state of the art and application to trauma mechanism classification

  • 融合表格与文本数据生成高质量合成医疗记录
  • 在创伤机制分类任务中实现高准确率,支持模型训练
  • 适合医疗AI研究者用于隐私保护下的算法开发

面对患者隐私保护和科研可复现性的挑战,医疗机器学习研究正转向合成医疗数据库的构建。本文简要综述了当前生成合成表格与文本数据的先进机器学习方法,重点聚焦其在创伤机制自动分类中的应用,并提出一种结合表格与非结构化文本数据的高质量合成医疗记录生成方法,旨在支持安全、高效、可复现的医疗人工智能研究。

原文摘要 · Abstract (English)

Faced with the challenges of patient confidentiality and scientific reproducibility, research on machine learning for health is turning towards the conception of synthetic medical databases. This article presents a brief overview of state-of-the-art machine learning methods for generating synthetic tabular and textual data, focusing their application to the automatic classification of trauma mechanisms, followed by our proposed methodology for generating high-quality, synthetic medical records combining tabular and unstructured text data.

合成数据医疗AI隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。