arXiv:2510.02734q-bio.BMcs.AI2025-10被引 1

用稀疏自编码器解析RNA语言模型的内部表征,发现与生物功能相关的特征。

SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations

  • 用稀疏自编码器分析RiNALMo模型的表征,提取可解释的特征。
  • 识别出与RNA家族身份和结构背景相关的稀疏特征成分。
  • 为比较不同RNA群体提供可解释的特征级分析框架。

深度学习,特别是大语言模型的发展,已推动生物分子建模的变革。以ESM为代表的蛋白质语言模型启发了新兴的RNA语言模型如RiNALMo。近期研究开始将稀疏自编码器(SAEs)应用于蛋白质语言模型表征,探索生物分子模型中的表征可解释性。本文探讨了SAEs能否为RNA语言模型表征提供可解释的特征分解,并考察其在此场景下的局限性。我们提出SAE-RNA,一种用于分析RiNALMo表征并映射到已知人类层面生物特征的可解释性模型。我们不宣称发现了确定的生物学概念,而是将基于SAE的分析视为对RNA语言模型内部信息组织方式的表征级探针。更广泛地说,SAE-RNA提供了一个特征级框架,可用于比较RNA群体,并识别与RNA家族身份或结构背景相关的稀疏表示成分。

原文摘要 · Abstract (English)

Deep learning, particularly with the advancement of Large Language Models, has transformed biomolecular modeling, with protein language models such as ESM inspiring emerging RNA language models such as RiNALMo. Recent work has begun applying sparse autoencoders (SAEs) to protein language model representations, exploring representation-level interpretability in biomolecular models. Here, we explore whether SAEs can provide interpretable feature decompositions of RNA language model representations, while also examining their limitations in this setting. We present SAE-RNA, interpretability model that analyzes RiNALMo representations and maps them to known human-level biological features. Rather than claiming definitive biological concept discovery, our study frames SAE-based analysis as a representation-level probe for characterizing how RNA language models organize biological information internally. More broadly, SAE-RNA provides a feature-level framework for comparing RNA groups and identifying sparse representation components associated with RNA family identity or structural context.

RNA语言模型稀疏自编码器可解释性特征分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。