新方法在无参考物种时仍能零误报筛查危险基因序列
CRC-Screen: Certified DNA-Synthesis Hazard Screening Under Taxonomic Shift

- 融合三类信号,用校准模型保证风险率上限
- 十次隔离分类法测试中零漏检、九次零误报
- 关键瓶颈是校准数据量,非算法本身
DNA合成服务商通过比对订单序列与预设危险列表进行筛查。我们发现,当危险序列来自参考集缺失的分类家族时,该基线方法的误报率高达100%:在符合性风险控制的认证漏检率约束下,低区分度信号导致阈值低于全部良性样本。本文构建三种源自公开注释的信号:k-mer Jaccard相似度、五模型评委组的修剪均值评分、聚类嵌入中心的余弦相似度。三者经单调逻辑聚合器融合,并由符合性风险控制校准。所得筛查器满足期望漏检率 ≤ α + TV,其中附加项为分类族留出下的分布偏移(跨折次认证上限24%-49%)。在UniProt KW-0800 reviewed毒素数据集上,α=0.05条件下,十次留一分类族外推测试中,所有折次实现0%实测漏检率,九折实现0%误报率。有限样本松弛项1/(n_cal+1)将可认证漏检率上限定为1.77%(基于200个危险序列子集);达到采购级α=10⁻³需校准集扩大18倍,全量已审核UniProt KW-0800语料库足以支撑。认证化筛查的决定性瓶颈在于校准数据,而非算法。
原文摘要 · Abstract (English)
DNA-synthesis providers screen incoming orders by searching the requested sequence against curated hazard lists. We show that this baseline collapses to a 100% false-flag rate when the hazardous sequence comes from a taxonomic family absent from the reference set: under Conformal Risk Control's certified miss-rate constraint, a low-discrimination signal forces the threshold below the entire test-benign mass. We compose three signals derived from a synthesis order's public annotation: $k$-mer Jaccard similarity to known toxins, the trimmed-mean score of a five-LLM judge panel, and cosine similarity to clustered embedding centroids. Fused under a monotone logistic aggregator and calibrated by Conformal Risk Control, the resulting screener certifies $\mathbb{E}[\mathrm{FNR}] \le α+ \mathrm{TV}$, where the additive term is the calibration-to-test distribution shift under family holdout (a certified ceiling of 24-49% across folds). Across ten leave-one-taxonomic-family-out folds at $α=0.05$ on UniProt KW-0800 reviewed toxins, the calibrated screener achieves 0% empirical test miss rate on every fold and 0% test false-flag rate on nine of ten folds. The bound's finite-sample slack $1/(n_{\mathrm{cal}}+1)$ caps the certifiable miss rate at 1.77% on our 200-hazard subsample; reaching procurement-grade $α=10^{-3}$ requires an $18\times$ larger calibration set, which the full reviewed UniProt KW-0800 corpus is large enough to deliver. The binding constraint on certifiable DNA-synthesis screening is calibration data, not algorithms. Code: https://github.com/najmulhasan-code/crc-screen
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。