arXiv:2504.07065q-bio.GNcs.LG2025-04中稿 · ICLR

边测序边分类,提升微生物鉴定速度与精度。

Enhancing Downstream Analysis in Genome Sequencing: Species Classification While Basecalling

  • 用多目标神经网络同步完成测序和物种分类。
  • 在17种细菌数据集上,顶1/3分类准确率达92.5%/98.9%。
  • 支持灵活配置,可选高精度或高速度模式。

快速准确地识别样本中的微生物物种(即宏基因组分析)在医疗和环境科学中至关重要。本文提出一种新方法,在测序设备产生信号的同时进行核酸序列解析(即测序)与多类基因组分类,采用多目标深度神经网络实现同步处理。引入新型损失策略:测序与分类的损失分别反向传播,共享层参数统一更新;并设计预设排名策略,支持用户选择前K个物种的识别精度,灵活权衡准确性与速度。实验结果表明,该方法在测序精度上达到当前最优水平,分类精度超过现有二分类模型,在包含17个基因组的Wick细菌数据集上,顶1/3物种识别平均准确率分别达92.5%/98.9%。该研究为宏基因组分析加速了关键瓶颈环节——将DNA序列匹配到正确基因组的过程。

原文摘要 · Abstract (English)

The ability to quickly and accurately identify microbial species in a sample, known as metagenomic profiling, is critical across various fields, from healthcare to environmental science. This paper introduces a novel method to profile signals coming from sequencing devices in parallel with determining their nucleotide sequences, a process known as basecalling, via a multi-objective deep neural network for simultaneous basecalling and multi-class genome classification. We introduce a new loss strategy where losses for basecalling and classification are back-propagated separately, with model weights combined for the shared layers, and a pre-configured ranking strategy allowing top-K species accuracy, giving users flexibility to choose between higher accuracy or higher speed at identifying the species. We achieve state-of-the-art basecalling accuracies, while classification accuracies meet and exceed the results of state-of-the-art binary classifiers, attaining an average of 92.5%/98.9% accuracy at identifying the top-1/3 species among a total of 17 genomes in the Wick bacterial dataset. The work presented here has implications for future studies in metagenomic profiling by accelerating the bottleneck step of matching the DNA sequence to the correct genome.

测序分类多任务学习宏基因组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。