用纳米孔测序原始信号直接聚类,提升DNA存储的准确性和速度。
Beyond the Alphabet: Deep Signal Embedding for Enhanced DNA Clustering
- 在碱基识别前,直接对纳米孔原始信号进行深度嵌入聚类。
- 准确率更高,计算时间减少30%以上,优于传统先碱基识别再聚类的方法。
- 适合从事高密度生物存储与测序算法优化的研究者。
DNA存储利用核酸碱基(A/T/C/G)作为数字信息的存储介质,具有超高密度和长期稳定性。其流程包括:(1) 将原始数据编码为DNA序列;(2) 合成这些序列并以无序集合形式储存;(3) 测序生成DNA读段;(4) 还原原始数据。合成与测序阶段会产生多个易出错的重复读段,用于最终重建原始序列。关键步骤是将读段聚类为源自同一原始链的组,再据此估计原始序列。现有方法在碱基识别(basecalling)后才进行聚类,但本工作提出在碱基识别前,直接利用纳米孔测序仪产生的原始信号进行深度神经网络嵌入聚类。该方法无需先转换为离散碱基,显著提升了聚类精度,并降低了计算开销,实验显示在相同数据集上比传统方法准确率更高,推理时间减少超过30%。
原文摘要 · Abstract (English)
The emerging field of DNA storage employs strands of DNA bases (A/T/C/G) as a storage medium for digital information to enable massive density and durability. The DNA storage pipeline includes: (1) encoding the raw data into sequences of DNA bases; (2) synthesizing the sequences as DNA \textit{strands} that are stored over time as an unordered set; (3) sequencing the DNA strands to generate DNA \textit{reads}; and (4) deducing the original data. The DNA synthesis and sequencing stages each generate several independent error-prone duplicates of each strand which are then utilized in the final stage to reconstruct the best estimate for the original strand. Specifically, the reads are first \textit{clustered} into groups likely originating from the same strand (based on their similarity to each other), and then each group approximates the strand that led to the reads of that group. This work improves the DNA clustering stage by embedding it as part of the DNA sequencing. Traditional DNA storage solutions begin after the DNA sequencing process generates discrete DNA reads (A/T/C/G), yet we identify that there is untapped potential in using the raw signals generated by the Nanopore DNA sequencing machine before they are discretized into bases, a process known as \textit{basecalling}, which is done using a deep neural network. We propose a deep neural network that clusters these signals directly, demonstrating superior accuracy, and reduced computation times compared to current approaches that cluster after basecalling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。