首个面向电磁领域的多模态基础模型,实现感知-识别-决策闭环
PReD: An LLM-based Foundation Multimodal Model for Electromagnetic Perception, Recognition, and Decision
- 构建多视角电磁数据集PReD-1.3M,覆盖波形、频谱等多表征
- 在6类任务上达到当前最优性能,支持端到端语言驱动决策
- 适合雷达通信、电子对抗领域研究者,推动智能电磁系统发展
多模态大模型在通用领域展现出强大的跨模态理解与推理能力,但在电磁(EM)领域仍面临数据稀缺和领域知识融合不足的挑战。本文提出PReD,首个覆盖‘感知-识别-决策’全链条的电磁领域基础模型。构建了高质量多任务电磁数据集PReD-1.3M与评估基准PReD-Bench,包含原始时域波形、频域谱图、星座图等多种表征,涵盖通信与雷达信号典型特征。支持信号检测、调制识别、参数估计、协议识别、射频指纹识别及抗干扰决策等核心任务。PReD采用多阶段训练策略,统一多个任务进行联合优化,实现从信号理解到语言驱动推理与决策的闭环,显著提升电磁领域专业能力,同时保持通用多模态能力。实验表明,PReD在基于开源与自采信号数据构建的PReD-Bench上取得领先性能,验证了视觉对齐基础模型在提升电磁信号理解与推理方面的可行性与潜力。
原文摘要 · Abstract (English)
Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, they still face challenges such as data scarcity and insufficient integration of domain knowledge. This paper proposes PReD, the first foundation model for the EM domain that covers the intelligent closed-loop of "perception, recognition, decision-making." We constructed a high-quality multitask EM dataset, PReD-1.3M, and an evaluation benchmark, PReD-Bench. The dataset encompasses multi-perspective representations such as raw time-domain waveform, frequency-domain spectrograms, and constellation diagrams, covering typical features of communication and radar signals. It supports a range of core tasks, including signal detection, modulation recognition, parameter estimation, protocol recognition, radio frequency fingerprint recognition, and anti-jamming decision-making. PReD adopts a multi-stage training strategy that unifies multiple tasks for EM signals. It achieves closed-loop optimization from end-to-end signal understanding to language-driven reasoning and decision-making, significantly enhancing EM domain expertise while maintaining general multimodal capabilities. Experimental results show that PReD achieves state-of-the-art performance on PReD-Bench constructed from both open-source and self-collected signal datasets. These results collectively validate the feasibility and potential of vision-aligned foundation models in advancing the understanding and reasoning of EM signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。