arXiv:2512.17527cs.LGcs.AI2025-12

基于物理化学特征的蛋白危险性筛查基准,支持CPU运行且防泄露。

SafeBench-Seq: A Homology-Clustered, CPU-Only Baseline for Protein Hazard Screening with Physicochemical/Composition Features and Cluster-Aware Confidence Intervals

  • 用氨基酸组成和理化特征构建可解释的分类器,仅依赖公开数据。
  • 在≤40%同源性聚类下评估,避免过拟合,提升真实威胁检测能力。
  • 提供校准概率与置信区间,适合安全研究与模型鲁棒性测试。

蛋白质设计的基础模型带来实际生物安全风险,但社区缺乏在同源性控制下、可在普通CPU上运行的可复现序列级危险筛查基线。我们提出SafeBench-Seq,一个仅依赖公开数据(SafeProtein危害序列与UniProt良性序列)和可解释特征(全局理化描述符与氨基酸组成)的元数据级基准与分类器。为模拟“从未见过”的威胁,将合并数据集按≤40%序列相似性进行同源聚类,并采用聚类级留出验证(训练集与测试集无聚类重叠)。报告区分度(AUROC/AUPRC)与筛查操作点(TPR@1% FPR;FPR@95% TPR),并计算95%自助法置信区间(n=200)。通过CalibratedClassifierCV提供校准概率(逻辑回归/随机森林用等熵校准;线性SVM用Platt sigmoid)。使用布里尔得分、期望校准误差(ECE,15分箱)与可靠性图评估概率质量。通过保留组成特征的残基置换与长度/组成仅有的消融实验探测捷径依赖。实证表明,随机划分显著高估模型鲁棒性;校准后的线性模型表现良好,而树集成模型略高布里尔得分与ECE。SafeBench-Seq仅发布元数据(访问编号、聚类ID、划分标签),不分发危险序列,实现可复现且安全的评估。

原文摘要 · Abstract (English)

Foundation models for protein design raise concrete biosecurity risks, yet the community lacks a simple, reproducible baseline for sequence-level hazard screening that is explicitly evaluated under homology control and runs on commodity CPUs. We introduce SafeBench-Seq, a metadata-only, reproducible benchmark and baseline classifier built entirely from public data (SafeProtein hazards and UniProt benigns) and interpretable features (global physicochemical descriptors and amino-acid composition). To approximate "never-before-seen" threats, we homology-cluster the combined dataset at <=40% identity and perform cluster-level holdouts (no cluster overlap between train/test). We report discrimination (AUROC/AUPRC) and screening-operating points (TPR@1% FPR; FPR@95% TPR) with 95% bootstrap confidence intervals (n=200), and we provide calibrated probabilities via CalibratedClassifierCV (isotonic for Logistic Regression / Random Forest; Platt sigmoid for Linear SVM). We quantify probability quality using Brier score, Expected Calibration Error (ECE; 15 bins), and reliability diagrams. Shortcut susceptibility is probed via composition-preserving residue shuffles and length-/composition-only ablations. Empirically, random splits substantially overestimate robustness relative to homology-clustered evaluation; calibrated linear models exhibit comparatively good calibration, while tree ensembles retain slightly higher Brier/ECE. SafeBench-Seq is CPU-only, reproducible, and releases metadata only (accessions, cluster IDs, split labels), enabling rigorous evaluation without distributing hazardous sequences.

蛋白安全可解释建模基准测试同源聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。