用大模型动态生成可解释的DNA特征,让决策树更准更易懂。
Interpretable DNA Sequence Classification via Dynamic Feature Generation in Decision Trees

- 在树分裂时动态生成生物相关序列特征,替代原始单核苷酸
- 在多个基因组任务中实现高预测性能且特征可人工解读
- 适合需要可解释性的基因功能分析与疾病机制研究
DNA序列分析在进化生物学、基因调控和疾病机制研究中至关重要。尽管深度神经网络预测效果优异,但其黑箱特性限制了理解。相比之下,轴对齐决策树具备可解释性,但因仅考虑单一原始特征导致表达能力受限,需极深树结构,影响可解释性和泛化性能。本文提出DEFT框架,在树构建过程中自适应生成高层次序列特征。DEFT利用大语言模型基于节点局部序列分布提出生物学合理的特征,并通过反思机制迭代优化。实验表明,DEFT在多种基因组任务中发现人类可读且高度预测的序列特征。
原文摘要 · Abstract (English)
The analysis of DNA sequences has become critical in numerous fields, from evolutionary biology to understanding gene regulation and disease mechanisms. While deep neural networks can achieve remarkable predictive performance, they typically operate as black boxes. Contrasting these black boxes, axis-aligned decision trees offer a promising direction for interpretable DNA sequence analysis, yet they suffer from a fundamental limitation: considering individual raw features in isolation at each split limits their expressivity, which results in prohibitive tree depths that hinder both interpretability and generalization performance. We address this challenge by introducing DEFT, a novel framework that adaptively generates high-level sequence features during tree construction. DEFT leverages large language models to propose biologically-informed features tailored to the local sequence distributions at each node and to iteratively refine them with a reflection mechanism. Empirically, we demonstrate that DEFT discovers human-interpretable and highly predictive sequence features across a diverse range of genomic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。