arXiv:2509.08300cs.LGcs.AI2025-09

用少量数据训练高精度调制识别模型,大幅降低训练成本。

\emph{FoQuS}: A Forgetting-Quality Coreset Selection Framework for Automatic Modulation Recognition

  • 根据样本训练过程中的预测变化,动态筛选关键数据子集。
  • 仅用1%-30%数据即可保持高识别准确率和跨模型泛化能力。
  • 适合需要快速迭代的通信信号识别系统研发人员。

基于深度学习的自动调制识别(AMR)模型在大规模标注数据支持下取得了显著进展。然而,在开发新模型或进行超参数调优时,重复使用海量数据训练带来的耗时与能耗往往难以承受。为此,我们提出FoQuS,通过从原始数据集中选择一个核心子集(coreset),近似全量数据训练效果,从而显著降低训练开销。具体而言,FoQuS记录每个样本在全数据集训练过程中的预测轨迹,并基于训练动态构建三个重要性度量。实验表明,使用仅1%-30%的原始数据,FoQuS可在多个AMR数据集上保持高识别准确率,并具备良好的跨架构泛化能力。

原文摘要 · Abstract (English)

Deep learning-based Automatic Modulation Recognition (AMR) model has made significant progress with the support of large-scale labeled data. However, when developing new models or performing hyperparameter tuning, the time and energy consumption associated with repeated training using massive amounts of data are often unbearable. To address the above challenges, we propose \emph{FoQuS}, which approximates the effect of full training by selecting a coreset from the original dataset, thereby significantly reducing training overhead. Specifically, \emph{FoQuS} records the prediction trajectory of each sample during full-dataset training and constructs three importance metrics based on training dynamics. Experiments show that \emph{FoQuS} can maintain high recognition accuracy and good cross-architecture generalization on multiple AMR datasets using only 1\%-30\% of the original data.

调制识别核心子集高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。