开源工具seqme统一评估生物序列设计质量,支持多类分子与多种评价指标。
seqme: a Python library for evaluating biological sequence design
- 模块化设计,支持序列、嵌入和性质三类评估指标
- 覆盖小分子到蛋白质全尺度生物序列,支持一次性与迭代设计方法
- 内置多种嵌入模型与可视化功能,适合算法开发与验证
近年来,计算方法在生物序列设计上的进展催生了多种评估指标,用于衡量生成序列对目标分布的保真度及所需特性的达成情况。然而,缺乏一个集成这些指标的单一软件库。本文提出seqme,一个模块化且高度可扩展的开源Python库,包含适用于生物序列设计方法的模型无关评估指标。seqme涵盖三类指标:基于序列、基于嵌入和基于性质,并适用于多种生物序列类型,包括小分子、DNA、ncRNA、mRNA、肽和蛋白质。该库提供多种生物序列的嵌入与性质模型,以及诊断与可视化函数,用于结果分析。seqme可用于评估一次性与迭代式计算设计方法。
原文摘要 · Abstract (English)
Recent advances in computational methods for designing biological sequences have sparked the development of metrics to evaluate these methods performance in terms of the fidelity of the designed sequences to a target distribution and their attainment of desired properties. However, a single software library implementing these metrics was lacking. In this work we introduce seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design. seqme considers three groups of metrics: sequence-based, embedding-based, and property-based, and is applicable to a wide range of biological sequences: small molecules, DNA, ncRNA, mRNA, peptides and proteins. The library offers a number of embedding and property models for biological sequences, as well as diagnostics and visualization functions to inspect the results. seqme can be used to evaluate both one-shot and iterative computational design methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。