首次为变压器信道解码器提供理论泛化保障,揭示其性能与模型规模、数据量的关系。
Generalization Bounds for Transformer Channel Decoders
- 通过比特级Rademacher复杂度建立泛化误差上界,连接噪声估计与误码率
- 证明码长、参数量和训练集大小均影响泛化能力,多层结构同样适用
- 发现校验位掩码注意力带来稀疏性,显著提升泛化性能,适合通信领域研究者
Transformer信道解码器(如错误校正码Transformer,ECCT)在实际应用中表现出色,但其泛化行为缺乏理论解释。本文从学习理论角度分析ECCT的泛化性能,通过建立乘性噪声估计误差与误码率(BER)之间的联系,利用比特级Rademacher复杂度推导出泛化差距的上界。该上界揭示了码长、模型参数和训练集大小对泛化能力的影响,适用于单层和多层ECCT。进一步发现,基于校验位的掩码注意力机制可引入稀疏性,降低覆盖数,从而获得更紧的泛化界。据我们所知,这是首次为这类解码器提供理论泛化保证。
原文摘要 · Abstract (English)
Transformer channel decoders, such as the Error Correction Code Transformer (ECCT), have shown strong empirical performance in channel decoding, yet their generalization behavior remains theoretically unclear. This paper studies the generalization performance of ECCT from a learning-theoretic perspective. By establishing a connection between multiplicative noise estimation errors and bit-error-rate (BER), we derive an upper bound on the generalization gap via bit-wise Rademacher complexity. The resulting bound characterizes the dependence on code length, model parameters, and training set size, and applies to both single-layer and multi-layer ECCTs. We further show that parity-check-based masked attention induces sparsity that reduces the covering number, leading to a tighter generalization bound. To the best of our knowledge, this work provides the first theoretical generalization guarantees for this class of decoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。