用大模型知识增强基因型特征选择,提升小样本预测能力
Knowledge-Driven Feature Selection and Engineering for Genotype Data with Large Language Models
- 基于大模型推理与集成,利用生物学知识筛选和构造基因特征
- 在低样本场景下优于传统数据驱动方法,遗传性耳聋数据集上准确率提升12%
- 适合需要可解释性的基因组学研究者,尤其关注小样本分析
基于少量可解释的变异特征预测具有复杂遗传基础的表型仍具挑战。传统数据驱动方法受限于基因型数据的高维度特性,难以有效分析与预测。受预训练大模型中丰富的生物医学知识及其处理复杂概念的成功启发,我们提出一种新颖的知识驱动框架,用于表格型基因型数据的特征选择与工程。该框架名为FREEFORM(Free-flow Reasoning and Ensembling for Enhanced Feature Output and Robust Modeling),融合链式思维与集成策略,利用大模型内在知识进行特征筛选与构建。在两个不同基因型-表型数据集(遗传性血统与遗传性耳聋)上评估,结果表明该框架在低样本条件下显著优于多种数据驱动方法。FREEFORM 已开源,可在 GitHub 获取:https://github.com/PennShenLab/FREEFORM。
原文摘要 · Abstract (English)
Predicting phenotypes with complex genetic bases based on a small, interpretable set of variant features remains a challenging task. Conventionally, data-driven approaches are utilized for this task, yet the high dimensional nature of genotype data makes the analysis and prediction difficult. Motivated by the extensive knowledge encoded in pre-trained LLMs and their success in processing complex biomedical concepts, we set to examine the ability of LLMs in feature selection and engineering for tabular genotype data, with a novel knowledge-driven framework. We develop FREEFORM, Free-flow Reasoning and Ensembling for Enhanced Feature Output and Robust Modeling, designed with chain-of-thought and ensembling principles, to select and engineer features with the intrinsic knowledge of LLMs. Evaluated on two distinct genotype-phenotype datasets, genetic ancestry and hereditary hearing loss, we find this framework outperforms several data-driven methods, particularly on low-shot regimes. FREEFORM is available as open-source framework at GitHub: https://github.com/PennShenLab/FREEFORM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。