arXiv:2506.11182q-bio.GNcs.AI2025-06中稿 · ICML被引 2

用预训练模型和染色质可及性数据提升CRISPR-Cas12引导RNA活性预测

Multimodal Modeling of CRISPR-Cas12 Activity Using Foundation Models and Chromatin Accessibility Data

  • 用已有的转录组预训练模型提取特征,搭配轻量回归器预测
  • 结合染色质可及性数据后性能显著提升,超越传统基线
  • 无需领域特定预训练,适合基因编辑研究者使用

预测引导RNA(gRNA)活性对高效CRISPR-Cas12基因编辑至关重要,但受限于数据不足、PAM序列差异及对大规模训练的依赖。本文探究了仅基于转录组预训练的生物基础模型是否能在无领域特定预训练的情况下提升gRNA活性估计。通过将现有RNA基础模型的嵌入作为轻量回归器输入,结果显示性能显著优于传统基线。进一步整合染色质可及性数据以捕捉调控背景,进一步提升预测效果。结果表明,预训练基础模型与染色质可及性数据在gRNA活性预测中具有显著有效性。

原文摘要 · Abstract (English)

Predicting guide RNA (gRNA) activity is critical for effective CRISPR-Cas12 genome editing but remains challenging due to limited data, variation across protospacer adjacent motifs (PAMs-short sequence requirements for Cas binding), and reliance on large-scale training. We investigate whether pre-trained biological foundation model originally trained on transcriptomic data can improve gRNA activity estimation even without domain-specific pre-training. Using embeddings from existing RNA foundation model as input to lightweight regressor, we show substantial gains over traditional baselines. We also integrate chromatin accessibility data to capture regulatory context, improving performance further. Our results highlight the effectiveness of pre-trained foundation models and chromatin accessibility data for gRNA activity prediction.

基因编辑预训练模型染色质可及性CRISPR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。