arXiv:2502.21060cs.LGcs.IT2025-02

用Transformer提升DNA存储中多类型错误的纠错能力

VT-Former: Efffcient Transformer-based Decoder for Varshamov-Tenengolts Codes

  • 基于统计特征与符号嵌入的Transformer架构实现高效解码
  • 单错纠正接近100%准确率,多错场景帧/比特准确率均提升
  • 解码延迟更低,适合高吞吐量的生物存储系统

近年来,针对基于DNA的数据存储中插入、删除和替换(IDS)错误的纠错问题受到广泛关注。在各类IDS纠错码中,最初为单错误校正设计的Varshamov-Tenengolts(VT)码已成为核心研究方向。现有解码方法虽在单错误校正上精度较高,但难以适用于多错误场景。本文提出一种增强统计特征的Transformer-based VT解码器(VT-Former),利用符号与统计特征嵌入挖掘VT码潜在的多错误校正能力。实验表明,VT-Former在单错误校正任务中达到近100%准确率;在不同码长的多错误解码任务中,相比传统硬判决与软进软出算法,帧准确率和比特准确率均有提升。此外,基础模型解码延迟低于传统软解码器,本研究进一步优化架构以提升效率并降低计算开销。

原文摘要 · Abstract (English)

In recent years, widespread attention has been drawn to the challenge of correcting insertion, deletion, and substitution (IDS) errors in DNA-based data storage. Among various IDS-correcting codes, Varshamov-Tenengolts (VT) codes, originally designed for single-error correction, have been established as a central research focus. While existing decoding methods demonstrate high accuracy for single-error correction, they are typically not applicable to the correction of multiple IDS errors. In this work, the latent capability of VT codes for multiple-error correction is investigated through a statistic-enhanced Transformer-based VT decoder (VT-Former), utilizing both symbol and statistic feature embeddings. Experimental results demonstrate that VT-Former achieves nearly 100\% accuracy on correcting single errors. For multi-error decoding tasks across various codeword lengths, improvements in both frame accuracy and bit accuracy are observed, compared to conventional hard-decision and soft-in soft-out decoding algorithms. Furthermore, while lower decoding latency is exhibited by the base model compared to traditional soft decoders, the architecture is further optimized in this study to enhance decoding efficiency and reduce computational overhead.

DNA存储纠错码Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。