用显式领域信号生成合成数据,让小模型也能媲美大模型。
ELTEX: A Framework for Domain-Driven Synthetic Data Generation
- 通过提取领域特征显式指导数据生成,保持专业知识完整。
- 在区块链攻击分类上,20亿参数模型性能接近GPT-4o,算力需求更低。
- 提供谷歌表格插件,非技术人员也能使用。
我们提出高效大模型分词提取(ELTEX)框架,解决大模型领域专精难题,通过系统性提取并整合领域信号,实现合成数据生成中的知识完整性。与依赖隐式知识迁移的方法不同,ELTEX显式利用领域信息。在网络安全案例研究中,使用ELTEX增强的数据训练的Gemma-2B模型,在区块链网络攻击分类任务上达到与GPT-4o相当的性能,同时显著降低计算开销。我们还开发了谷歌表格插件,使非技术用户可轻松使用。贡献包括:(1) ELTEX框架;(2) 谷歌表格附加组件;(3) 实验验证其缩小小模型与大模型间性能差距的效果;(4) 一个包含11,448条文本的区块链攻击检测合成数据集。
原文摘要 · Abstract (English)
We introduce Efficient LLM Token Extraction (ELTEX), a framework addressing the critical challenge of LLM domain specialization by systematically extracting and integrating domain indicators throughout synthetic data generation. Unlike approaches relying on implicit knowledge transfer, ELTEX explicitly leverages domain signals to maintain specialized knowledge integrity. In our cybersecurity case study, ELTEX-enhanced data enables a fine-tuned Gemma-2B model to achieve performance competitive with GPT-4o on blockchain cyberattack classification while reducing computational requirements. Our Google Sheets implementation makes ELTEX accessible to non-technical users. Our contributions include: (1) the ELTEX framework; (2) Google Sheets Add-on implementation; (3) empirical validation showing how ELTEX bridges performance gaps between small and large models; and (4) a synthetic dataset of 11,448 texts for blockchain cyberattack detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。