arXiv:2606.06834cs.CLq-bio.GN2026-06

区分基因组预测性与调控性,揭示胶质瘤中真正起作用的非编码元件。

The Dark Regulome: Disentangling Predictability from Regulation in Genomic Foundation Models

论文配图:The Dark Regulome: Disentangling Predictability from Regulation in Genomic Foundation Models
图 1 · 摘自论文原文
  • 提出残差-置换诊断法,分离模型预测性与真实调控信号。
  • 发现10kb以内的近端调控区是核心,且六特征线性模型已达AUC 0.985。
  • 跨模型分析显示调控信号与序列可预测性几乎无重叠,适合生物机制研究者。

高级别胶质瘤通过功能突触整合进神经回路,引发关于哪些非编码元件调控肿瘤细胞突触生成基因表达的问题。我们称这些隐藏在基因组中的调控程序为‘暗调控组’。尽管序列基础模型可通过零样本体外诱变(ISM)探测,但基于似然的评分天然依赖局部序列可预测性,导致调控解释模糊。在三个架构不同的基础模型(Caduceus-Ph、HyenaDNA、Enformer)和30,448个位于92个胶质瘤相关位点的暗基因组元件上,我们引入残差-置换诊断法,将预测性驱动与调控驱动的RIS方差解耦。无论何种控制,10kb以内的近端调控边界均稳定存在,而语言模型推导的元件层级结构不具鲁棒性:仅六特征线性基线即可在Caduceus模型中实现前10%成员识别,AUC达0.985。跨架构分解清晰划分出两个层面:序列可预测层(两种语言模型共同排序高度可预测的转座元件),以及调控输出层(仅Enformer保有残差的cCRE判别信号),二者前100名列表几乎无重叠。进化保守性、脑部顺式-eQTL及STRING-PPI交叉验证确认:所有三模型前100名元件每模型均富集3.3倍脑eQTL(p_emp < 5×10⁻³);而看似诱人的转座元件调控层与显著的NRXN1+NLGN1蛋白对收敛现象,在严格置换测试下均被排除。本研究提供一种通用方法学工具,适用于所有基于ISM的调控研究。

原文摘要 · Abstract (English)

High-grade gliomas integrate into neural circuits through functional synapses with neurons, raising the question of which noncoding elements shape synaptogenic gene expression in tumor cells. The regulatory program written across the dark genome, what we call the $\textit{dark regulome}$, is the natural substrate to probe, and sequence foundation models offer a zero-shot route through in-silico mutagenesis (ISM); yet likelihood-based scoring is tautologically coupled to local sequence predictability, leaving the regulatory interpretation underdetermined. Across three architecturally distinct foundation models (Caduceus-Ph, HyenaDNA, Enformer) and 30,448 dark genome elements at 92 glioma-relevant loci, we introduce a residualization-and-permutation diagnostic that separates predictability-driven from regulation-driven RIS variance. A sharp 10kb proximal-regulatory horizon survives every control we apply, but the LM-derived element-class hierarchy does not: a six-feature linear baseline matches Caduceus top-decile membership at AUC $= 0.985$. Cross-architecture decomposition cleanly separates a sequence-predictability layer (the two language models co-rank long well-predicted transposable elements) from a regulatory-output layer (Enformer alone retains residual cCRE-discriminative signal), with literally zero overlap between the two top-100 lists. Conservation, brain cis-eQTL, and STRING-PPI cross-checks then anchor what biology survives: top-100 elements across all three models are $3.3\times$ enriched per model for matching brain eQTLs ($p_\mathrm{emp} < 5\times 10^{-3}$), while a tempting transposable-element regulatory layer and a striking NRXN1+NLGN1 protein-pair convergence both fail proper permutation tests once those tests are constructed. We deliver the diagnostic as a general methodological tool for any ISM-based regulatory study.

基因组学调控机制深度学习胶质瘤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。