arXiv:2507.16978q-bio.GNcs.AI2025-07被引 3

用向量搜索替代传统比对,加速基因序列相似性查找

Fast and Scalable Gene Embedding Search: A Comparative Study of FAISS and ScaNN

  • 对比FAISS与ScaNN在基因嵌入向量上的搜索性能
  • 内存与运行时间效率显著提升,检索准确率更高
  • 适合处理海量未注释基因数据的生物信息学研究

DNA测序数据的指数级增长已超出传统启发式方法的处理能力,这些方法在大规模场景下难以有效扩展。高效计算方法亟需用于支持大规模相似性搜索,这是生物信息学中检测同源性、功能相似性和序列新颖性的基础任务。尽管BLAST等工具仍广泛使用且在多数场景有效,但存在计算成本高、对差异序列表现差等问题。本文探索基于嵌入的相似性搜索方法,学习能捕捉深层结构与功能模式的潜在表示,超越原始序列比对。我们系统评估了两种最先进的向量搜索库——FAISS与ScaNN——在生物有意义的基因嵌入上的表现。不同于以往研究,本分析聚焦于生物信息学特定的嵌入,并检验其在发现新序列(如未表征类群或无已知同源物的基因)中的实用性。结果表明,该方法在内存与运行时效率方面具有显著优势,同时提升了检索质量,为传统依赖比对的工具提供了一种有前景的替代方案。

原文摘要 · Abstract (English)

The exponential growth of DNA sequencing data has outpaced traditional heuristic-based methods, which struggle to scale effectively. Efficient computational approaches are urgently needed to support large-scale similarity search, a foundational task in bioinformatics for detecting homology, functional similarity, and novelty among genomic and proteomic sequences. Although tools like BLAST have been widely used and remain effective in many scenarios, they suffer from limitations such as high computational cost and poor performance on divergent sequences. In this work, we explore embedding-based similarity search methods that learn latent representations capturing deeper structural and functional patterns beyond raw sequence alignment. We systematically evaluate two state-of-the-art vector search libraries, FAISS and ScaNN, on biologically meaningful gene embeddings. Unlike prior studies, our analysis focuses on bioinformatics-specific embeddings and benchmarks their utility for detecting novel sequences, including those from uncharacterized taxa or genes lacking known homologs. Our results highlight both computational advantages (in memory and runtime efficiency) and improved retrieval quality, offering a promising alternative to traditional alignment-heavy tools.

基因搜索向量检索FAISSScaNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。