自动化挖掘生物医学数据中潜在的有趣科学特征
InterFeat: A Pipeline for Finding Interesting Scientific Features
- 融合机器学习与大模型,定义并识别有趣性特征
- 在8种疾病中提前数年发现风险因子,验证率40%-53%
- 适合医学研究者快速发现新假说,提升科研效率
发现有趣现象是科学发现的核心,但这一概念长期依赖人工且定义模糊。本文提出一个整合式流水线,用于在结构化生物医学数据中自动发现有趣的简单假设(即特征与目标之间的效应方向及潜在机制)。该流水线结合机器学习、知识图谱、文献检索和大语言模型,将“有趣性”形式化为新颖性、实用性与可信度的综合指标。在英国生物银行涵盖的8种重大疾病数据上,该方法可稳定地在文献发表前数年发现风险因子。前若干候选中有40%–53%被验证为有趣,显著优于基于SHAP的基线(0%–7%)。整体上,109个候选中28%获得医学专家认可。该流水线实现了对“有趣性”的可扩展、通用化操作,支持任意目标。数据与代码已开源:https://github.com/LinialLab/InterFeat
原文摘要 · Abstract (English)
Finding interesting phenomena is the core of scientific discovery, but it is a manual, ill-defined concept. We present an integrative pipeline for automating the discovery of interesting simple hypotheses (feature-target relations with effect direction and a potential underlying mechanism) in structured biomedical data. The pipeline combines machine learning, knowledge graphs, literature search and Large Language Models. We formalize "interestingness" as a combination of novelty, utility and plausibility. On 8 major diseases from the UK Biobank, our pipeline consistently recovers risk factors years before their appearance in the literature. 40--53% of our top candidates were validated as interesting, compared to 0--7% for a SHAP-based baseline. Overall, 28% of 109 candidates were interesting to medical experts. The pipeline addresses the challenge of operationalizing "interestingness" scalably and for any target. We release data and code: https://github.com/LinialLab/InterFeat
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。