首个专用于通信系统设计的推理大模型,让AI自动完成6G系统构建。
DeepForm: Reasoning Large Language Model for Communication System Formulation
- 用思维链数据微调+规则强化学习,训练专用通信推理模型。
- 在多个场景中超越更大规模商用大模型,表现领先。
- 开源首个通信系统设计推理数据集,适合研究者和工程师使用。
通信系统建模对6G及未来无线技术发展至关重要,但仍是高度依赖专业知识的复杂任务。尽管大语言模型(LLM)具有潜力,现有通用模型普遍缺乏领域专知、精细推理能力以及高质量领域训练数据。为此,我们提出DeepForm,首个专用于自动化通信系统建模的推理大模型。我们构建了首个大规模、开源的领域专用数据集——通信系统建模推理语料库(CSFRC)。框架采用两阶段训练:首先通过带思维链(CoT)数据的监督微调(SFT)注入领域知识;其次引入基于ReMax改进的新型规则强化学习算法C-ReMax,培养高级建模能力,激发自我修正与验证等复杂推理模式。大量实验表明,该模型在多种场景下性能显著优于更大规模的私有化大模型。论文被接受后,我们将公开相关资源,推动该领域进一步研究。
原文摘要 · Abstract (English)
Communication system formulation is critical for advancing 6G and future wireless technologies, yet it remains a complex, expertise-intensive task. While Large Language Models (LLMs) offer potential, existing general-purpose models often lack the specialized domain knowledge, nuanced reasoning capabilities, and access to high-quality, domain-specific training data required for adapting a general LLM into an LLM specially for communication system formulation. To bridge this gap, we introduce DeepForm, the first reasoning LLM specially for automated communication system formulation. We propose the world-first large-scale, open-source dataset meticulously curated for this domain called Communication System Formulation Reasoning Corpus (CSFRC). Our framework employs a two-stage training strategy: first, Supervised Fine-Tuning (SFT) with Chain-of-Thought (CoT) data to distill domain knowledge; second, a novel rule-based Reinforcement Learning (RL) algorithm, C-ReMax based on ReMax, to cultivate advanced modeling capabilities and elicit sophisticated reasoning patterns like self-correction and verification. Extensive experiments demonstrate that our model achieves state-of-the-art performance, significantly outperforming larger proprietary LLMs on diverse senerios. We will release related resources to foster further research in this area after the paper is accepted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。