用生成模型造数据,解决数据少、隐私和标注难问题
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
- 用大语言模型、扩散模型等生成合成数据
- 提升数据挖掘中的数据可用性与隐私保护能力
- 适合研究数据稀缺场景的学者与工程师
大型语言模型、扩散模型和生成对抗网络等生成模型近期彻底改变了合成数据的创建方式,为数据挖掘中的数据稀缺、隐私保护和标注难题提供了可扩展的解决方案。本教程介绍合成数据生成的基础理论与最新进展,涵盖关键方法与实用框架,并讨论评估策略与应用场景。参与者将获得利用生成式合成数据提升数据挖掘研究与实践的切实洞见。更多信息请访问:https://syndata4dm.github.io/
原文摘要 · Abstract (English)
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalable solutions to data scarcity, privacy, and annotation challenges in data mining. This tutorial introduces the foundations and latest advances in synthetic data generation, covers key methodologies and practical frameworks, and discusses evaluation strategies and applications. Attendees will gain actionable insights into leveraging generative synthetic data to enhance data mining research and practice. More information can be found on our website: https://syndata4dm.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。