无需人工标注,AI从海量生物数据中自动发现疾病特征与基因关联。
Transcending the Annotation Bottleneck: AI-Powered Discovery in Biology and Medicine
- 通过无监督学习直接挖掘数据内在结构,跳过标注瓶颈。
- 在心脏性状遗传分析、组织切片基因表达预测上表现媲美有监督模型。
- 适合生物医学数据挖掘、临床辅助诊断等无需标签场景的研究者。
长期以来,专家标注限制了人工智能在生物医学中的应用。尽管监督学习推动了早期临床算法发展,但当前向无监督和自监督学习(SSL)的范式转变正释放大规模生物银行数据的潜力。这些方法直接从数据内在结构——如磁共振图像(MRI)的像素、三维扫描的体素或基因序列的词元——中学习,实现新表型发现、形态与遗传关联解析及异常检测,且不受人为偏见影响。本文综述了该领域的里程碑进展,展示了无监督框架如何推导可遗传的心脏性状,预测组织学中的空间基因表达,并以媲美或超越监督模型的性能检测病理性病变。
原文摘要 · Abstract (English)
The dependence on expert annotation has long constituted the primary rate-limiting step in the application of artificial intelligence to biomedicine. While supervised learning drove the initial wave of clinical algorithms, a paradigm shift towards unsupervised and self-supervised learning (SSL) is currently unlocking the latent potential of biobank-scale datasets. By learning directly from the intrinsic structure of data - whether pixels in a magnetic resonance image (MRI), voxels in a volumetric scan, or tokens in a genomic sequence - these methods facilitate the discovery of novel phenotypes, the linkage of morphology to genetics, and the detection of anomalies without human bias. This article synthesises seminal and recent advances in "learning without labels," highlighting how unsupervised frameworks can derive heritable cardiac traits, predict spatial gene expression in histology, and detect pathologies with performance that rivals or exceeds supervised counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。