文本数据蒸馏将大语料压缩为小量合成文本,提升训练效率。
Technical Report on Text Dataset Distillation
- 基于Transformer生成离散合成文本,模拟原始数据分布
- 已适配超10亿参数的解码器模型,支持复杂任务蒸馏
- 仍需统一评测标准,适合研究高效训练与数据压缩的学者
在视觉领域,数据蒸馏技术可将大规模数据集压缩为少量合成数据,在训练中保持相近效果。尽管图像数据蒸馏已有丰富成果,文本数据蒸馏仍处发展初期。早期方法借鉴视觉领域经验,但因文本的离散特性带来挑战,逐步演变为独立研究方向。近年来关键进展包括:利用Transformer模型生成离散合成文本、向超过10亿参数的解码器模型扩展应用。尽管现代方法取得显著进步,该领域仍处于成熟阶段,存在评测标准不统一、处理文本离散性困难、难以应对复杂任务及缺乏真实场景应用案例等问题。本文综述文本数据蒸馏的过往与最新进展,分析不同蒸馏策略、核心贡献与普遍挑战。
原文摘要 · Abstract (English)
In the vision domain, dataset distillation arises as a technique to condense a large dataset into a smaller synthetic one that exhibits a similar result in the training process. While image data presents an extensive literature of distillation methods, text dataset distillation has fewer works in comparison. Text dataset distillation initially grew as an adaptation of efforts from the vision universe, as the particularities of the modality became clear obstacles, it rose into a separate branch of research. Several milestones mark the development of this area, such as the introduction of methods that use transformer models, the generation of discrete synthetic text, and the scaling to decoder-only models with over 1B parameters. Despite major advances in modern approaches, the field remains in a maturing phase, with room for improvement on benchmarking standardization, approaches to overcome the discrete nature of text, handling complex tasks, and providing explicit examples of real-world applications. In this report, we review past and recent advances in dataset distillation for text, highlighting different distillation strategies, key contributions, and general challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。