协议感知分词是无线数据模型性能关键,架构仅影响部署效率。
Protocol-Aware Tokenization and Architecture Co-Design for Wireless Packet Foundation Models
- 采用协议感知分词,提升令牌表示能力
- 深度GPT达98.2%准确率,较原模型提升32点
- 分词器决定性能上限,架构适配部署需求
构建无线包迹基础模型的关键在于分词器、架构或二者协同?基于PLUME Anonymous[2026]提出的802.11分组协议感知分词方法,我们扩展模型深度并迁移同一分词器至不同架构。深度GPT(PLUME-DEEP,24层)达到98.2%的top-1准确率,相比原12层设计提升32个百分点;Mamba-2状态空间变体(PLUME-MAMBA)实现96.1%准确率,吞吐量提高1.7倍,上下文长度翻倍。通过控制变量的2×2对比发现:更换分词器导致准确率波动32点,而更换架构仅带来2点变化。核心结论:协议感知分词是性能主因,骨干网络则成为速度与精度权衡的部署调节器。
原文摘要 · Abstract (English)
What matters more for building foundation models for wireless packet traces: the tokenizer or the architecture or both? To answer this question, we build on PLUME Anonymous [2026], which introduced protocol-aware tokenization for 802.11 traces; we scale model depth and transfer the same tokenizer to a fundamentally different architecture family. A deeper GPT (PLUME-DEEP, 24 layers) reaches 98.2% top-1 accuracy, gaining 32 points over the original 12-layer design, while a Mamba-2 state-space variant (PLUME-MAMBA) achieves 96.1% with 1.7x higher throughput and 2x longer context. The key insight emerges from a controlled 2x2 comparison across tokenizers and architectures: changing the tokenizer swings accuracy by 32 points; changing the architecture moves it by only 2. Protocol-aware tokenization is the primary performance lever, and the backbone becomes a deployment knob trading accuracy for speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。