arXiv:2411.03522q-bio.GNcs.AI2024-11被引 2

用大模型分析长非编码RNA转录调控,发现效果好但受限于数据和任务复杂度。

Exploring the Potentials and Challenges of Using Large Language Models for the Analysis of Transcriptional Regulation of Long Non-coding RNAs

  • 用微调的基因组基础模型分析lncRNA序列调控
  • 模型在复杂任务上表现良好,但效果随任务难度下降
  • 适合研究基因调控机制的生物信息学工作者参考

长非编码RNA(lncRNA)在基因调控和疾病机制中起关键作用,但其序列复杂多变,功能机制与表达调控知识有限,给研究带来挑战。鉴于大语言模型(LLMs)在捕捉序列数据复杂依赖关系上的成功,本研究系统探索了LLMs在lncRNA基因转录调控序列分析中的潜力与局限。大量实验表明,微调后的基因组基础模型在逐步复杂的任务中展现出良好性能。此外,我们深入分析了任务复杂度、模型选择、数据质量及生物学可解释性对lncRNA表达调控研究的关键影响。

原文摘要 · Abstract (English)

Research on long non-coding RNAs (lncRNAs) has garnered significant attention due to their critical roles in gene regulation and disease mechanisms. However, the complexity and diversity of lncRNA sequences, along with the limited knowledge of their functional mechanisms and the regulation of their expressions, pose significant challenges to lncRNA studies. Given the tremendous success of large language models (LLMs) in capturing complex dependencies in sequential data, this study aims to systematically explore the potential and limitations of LLMs in the sequence analysis related to the transcriptional regulation of lncRNA genes. Our extensive experiments demonstrated promising performance of fine-tuned genome foundation models on progressively complex tasks. Furthermore, we conducted an insightful analysis of the critical impact of task complexity, model selection, data quality, and biological interpretability for the studies of the regulation of lncRNA gene expression.

lncRNA大模型基因调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。