arXiv:2508.20193cs.CVeess.SP2025-08被引 2

用重建驱动的视觉变压器,少标签也能精准识别信号调制方式。

Enhancing Automatic Modulation Recognition With a Reconstruction-Driven Vision Transformer Under Limited Labels

  • 融合监督、自监督和重建目标,统一训练框架提升特征学习。
  • 仅用15%-20%标签即达到ResNet级准确率,低标签下性能领先。
  • 适合标签稀缺场景,如频谱监测与认知无线电系统部署。

自动调制识别(AMR)在认知无线电、频谱监控和无线安全通信中至关重要。现有方法通常依赖大规模标注数据或多阶段训练流程,限制了实际应用中的可扩展性与泛化能力。本文提出一种统一的视觉变压器(ViT)框架,整合监督、自监督与重建目标。模型包含一个ViT编码器、轻量卷积解码器和线性分类器;重建分支将增强信号映射回原始形式,使编码器锚定于精细的I/Q结构。该策略在预训练中促进鲁棒且具有判别力的特征学习,而微调阶段部分标签监督则实现少量标签下的有效分类。在RML2018.01A数据集上,本方法在低标签条件下优于监督型CNN与ViT基线,仅需15%-20%标签即可接近ResNet准确率,并在不同信噪比(SNR)下保持优异性能。整体框架提供了一种简单、通用且标签高效的AMR解决方案。

原文摘要 · Abstract (English)

Automatic modulation recognition (AMR) is critical for cognitive radio, spectrum monitoring, and secure wireless communication. However, existing solutions often rely on large labeled datasets or multi-stage training pipelines, which limit scalability and generalization in practice. We propose a unified Vision Transformer (ViT) framework that integrates supervised, self-supervised, and reconstruction objectives. The model combines a ViT encoder, a lightweight convolutional decoder, and a linear classifier; the reconstruction branch maps augmented signals back to their originals, anchoring the encoder to fine-grained I/Q structure. This strategy promotes robust, discriminative feature learning during pretraining, while partial label supervision in fine-tuning enables effective classification with limited labels. On the RML2018.01A dataset, our approach outperforms supervised CNN and ViT baselines in low-label regimes, approaches ResNet-level accuracy with only 15-20% labeled data, and maintains strong performance across varying SNR levels. Overall, the framework provides a simple, generalizable, and label-efficient solution for AMR.

调制识别视觉变压器少样本学习信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。