用大模型生成无线数据,少样本下也能高效提升模型性能
LLM-AUG: Robust Wireless Data Augmentation with In-Context Learning in Large Language Models

- 通过提示工程在大模型中直接生成无线信号数据
- 仅用15%标注数据就接近理想性能,低信噪比下提升29.4%
- 适合资源受限、数据难获取的无线通信场景
数据稀缺是深度学习应用于无线通信的核心瓶颈,尤其在射频(RF)数据标注成本高、耗时长或受操作限制的情况下。本文提出LLM-AUG框架,利用大语言模型的上下文学习能力,在学习到的嵌入空间中生成合成训练样本。不同于需训练专用生成模型的传统方法,LLM-AUG通过结构化提示实现快速适应,适用于小样本场景。我们在RadioML 2016.10A和干扰分类(IC)数据集上评估了该方法,结果表明在低样本设置下,LLM-AUG始终优于传统增强与深度生成基线,并在仅使用15%标注数据时达到接近理论最优性能。在分布偏移下表现更鲁棒,于低信噪比条件下相较扩散模型提升29.4%;在RadioML与IC数据集上分别取得67.6%和35.7%的相对增益。t-SNE可视化显示合成样本有效保留类别结构,提升增广的一致性与信息量。结果证明,大语言模型可作为高效实用的无线机器学习数据增强器,支持在动态无线环境中实现稳健且数据高效的建模。
原文摘要 · Abstract (English)
Data scarcity remains a fundamental bottleneck in applying deep learning to wireless communication problems, particularly in scenarios where collecting labeled Radio Frequency (RF) data is expensive, time-consuming, or operationally constrained. This paper proposes LLM-AUG, a data augmentation framework that leverages in-context learning in large language models (LLMs) to generate synthetic training samples directly in a learned embedding space. Unlike conventional generative approaches that require training task-specific models, LLM-AUG performs data generation through structured prompting, enabling rapid adaptation in low-shot regimes. We evaluate LLM-AUG on two representative tasks: modulation classification and interference classification using the RadioML 2016.10A dataset, and the Interference Classification (IC) dataset respectively. Results show that LLM-AUG consistently outperforms traditional augmentation and deep generative baselines across low-shot settings and reaches near oracle performance using only 15% labeled data. LLM-AUG further demonstrates improved robustness under distribution shifts, yielding a 29.4% relative gain over diffusion-based augmentation at a lower SNR value. On the RadioML and IC datasets, LLM-AUG yields a relative gain of 67.6% and 35.7% over the diffusion-based baseline. The t-SNE visualizations further validate that synthetic samples generated by better preserve class structure in the embedding space, leading to more consistent and informative augmentations. These results demonstrate that LLMs can serve as effective and practical data augmenters for wireless machine learning, enabling robust and data-efficient learning in evolving wireless environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。