用深度学习提升DNA存储的纠错能力,实现高保真长期数据存取。
NEURODNAAI: Neural pipeline approaches for the advancing dna-based information storage as a sustainable digital medium using deep learning framework
- 融合生物约束与深度学习,构建端到端编码重建管道。
- 在文本和图像数据上实现极低比特错误率,优于传统方法。
- 适合需要超长期、高密度存档的科研与机构用户。
DNA是极具前景的数字信息存储介质,具备超高密度和长久稳定性。尽管已有研究推进了编码理论、工作流设计与仿真工具,但合成成本高、测序错误及生物限制(如GC含量失衡、同聚物)仍制约实际应用。为此,本框架借鉴量子并行思想,增强编码多样性与鲁棒性,将生物约束融入深度学习,以提升DNA存储中的错误缓解能力。NeuroDNAAI将二进制数据流编码为符号化DNA序列,经含替换、插入、删除的噪声信道传输后,可高保真重建。实验表明,传统提示或规则方案难以适应真实噪声,而NeuroDNAAI表现显著更优。在基准数据集上的测试显示,文本与图像数据均实现低比特错误率。通过统一理论、流程与仿真,该框架实现了可扩展、生物学合规的归档式DNA存储。
原文摘要 · Abstract (English)
DNA is a promising medium for digital information storage for its exceptional density and durability. While prior studies advanced coding theory, workflow design, and simulation tools, challenges such as synthesis costs, sequencing errors, and biological constraints (GC-content imbalance, homopolymers) limit practical deployment. To address this, our framework draws from quantum parallelism concepts to enhance encoding diversity and resilience, integrating biologically informed constraints with deep learning to enhance error mitigation in DNA storage. NeuroDNAAI encodes binary data streams into symbolic DNA sequences, transmits them through a noisy channel with substitutions, insertions, and deletions, and reconstructs them with high fidelity. Our results show that traditional prompting or rule-based schemes fail to adapt effectively to realistic noise, whereas NeuroDNAAI achieves superior accuracy. Experiments on benchmark datasets demonstrate low bit error rates for both text and images. By unifying theory, workflow, and simulation into one pipeline, NeuroDNAAI enables scalable, biologically valid archival DNA storage
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。