用大模型自动生成敏感领域高质量数据,解决标注难、数据少问题
FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
- 基于大模型动态生成语义连贯的领域数据,支持隐私保护场景
- 四组件协同生成,确保数据多样性和准确性,提升模型训练效果
- 适合医疗、安全等数据受限领域的研究人员使用
数据可用性与质量是机器学习中的关键挑战,尤其在数据稀缺、获取成本高或受隐私法规限制的领域。医疗、生物医学研究和网络安全等领域常面临高昂的数据获取成本、标注数据有限以及关键事件稀少或敏感等问题。这些因素共同构成‘数据挑战’,阻碍了高风险领域中准确且泛化能力强的机器学习模型发展。为此,我们提出 FlexiDataGen,一种用于敏感领域动态语义数据生成的大语言模型框架。该框架可自主合成丰富、语义一致且语言多样的领域数据。其整合四个核心组件:(1) 句法-语义分析,(2) 检索增强生成,(3) 动态元素注入,(4) 迭代改写与语义验证。这四者协同工作,保障生成数据的高质量与领域相关性。实验表明,FlexiDataGen有效缓解数据短缺与标注瓶颈,支持可扩展、高精度的机器学习模型开发。
原文摘要 · Abstract (English)
Dataset availability and quality remain critical challenges in machine learning, especially in domains where data are scarce, expensive to acquire, or constrained by privacy regulations. Fields such as healthcare, biomedical research, and cybersecurity frequently encounter high data acquisition costs, limited access to annotated data, and the rarity or sensitivity of key events. These issues-collectively referred to as the dataset challenge-hinder the development of accurate and generalizable machine learning models in such high-stakes domains. To address this, we introduce FlexiDataGen, an adaptive large language model (LLM) framework designed for dynamic semantic dataset generation in sensitive domains. FlexiDataGen autonomously synthesizes rich, semantically coherent, and linguistically diverse datasets tailored to specialized fields. The framework integrates four core components: (1) syntactic-semantic analysis, (2) retrieval-augmented generation, (3) dynamic element injection, and (4) iterative paraphrasing with semantic validation. Together, these components ensure the generation of high-quality, domain-relevant data. Experimental results show that FlexiDataGen effectively alleviates data shortages and annotation bottlenecks, enabling scalable and accurate machine learning model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。