新基准揭示RNA结构预测模型泛化能力有限,仅部分模型在分布外仍稳健。
Fair splits flip the leaderboard: CHANRG reveals limited generalization in RNA secondary-structure prediction
- 构建17万条非冗余RNA数据集,按基因组分布设计划分方式
- 基础模型在测试集表现最佳,但跨家族预测时性能大幅下降
- 提出无填充、对称感知评估框架,提升模型鲁棒性验证可信度
准确预测RNA二级结构对转录组注释、非编码RNA机制分析及RNA药物设计至关重要。近期深度学习与RNA基础模型虽取得进展,但现有基准可能高估其跨RNA家族的泛化能力。本文提出综合层次化非编码RNA分组标注(CHANRG),基于超过1000万条Rfam 15.0序列,通过结构感知去重、基因组感知划分和多尺度结构评估,构建包含170,083条结构非冗余RNA的数据集。在29种预测器上对比发现,基础模型在留出数据上准确率最高,但在分布外表现显著退化;而结构解码器与直接神经网络预测器则保持更强鲁棒性。该差距在控制序列长度后依然存在,反映结构性覆盖缺失与高级连接错误。结合无填充、对称感知评估栈,CHANRG为开发具备真实分布外鲁棒性的RNA结构预测模型提供了更严格、批次无关的评估框架。
原文摘要 · Abstract (English)
Accurate prediction of RNA secondary structure underpins transcriptome annotation, mechanistic analysis of non-coding RNAs, and RNA therapeutic design. Recent gains from deep learning and RNA foundation models are difficult to interpret because current benchmarks may overestimate generalization across RNA families. We present the Comprehensive Hierarchical Annotation of Non-coding RNA Groups (CHANRG), a benchmark of 170{,}083 structurally non-redundant RNAs curated from more than 10 million sequences in Rfam~15.0 using structure-aware deduplication, genome-aware split design and multiscale structural evaluation. Across 29 predictors, foundation-model methods achieved the highest held-out accuracy but lost most of that advantage out of distribution, whereas structured decoders and direct neural predictors remained markedly more robust. This gap persisted after controlling for sequence length and reflected both loss of structural coverage and incorrect higher-order wiring. Together, CHANRG and a padding-free, symmetry-aware evaluation stack provide a stricter and batch-invariant framework for developing RNA structure predictors with demonstrable out-of-distribution robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。