用可解释的基因组模型,让机器人自动判断艾滋病病毒类型
Automated Genomic Interpretation via Concept Bottleneck Models for Medical Robotics
- 将DNA序列转为生物有意义的概念,再做分类决策
- 在自建和LANL数据集上准确区分艾滋病毒亚型
- 输出可验证的解释,适合临床自动化系统使用
我们提出一种自动化基因组解读模块,将原始DNA序列转化为可行动、可解释的决策,适用于医疗自动化与机器人系统。该框架结合混沌游戏表示(CGR)与概念瓶颈模型(CBM),强制预测通过GC含量、CpG密度、kmer基序等生物学意义明确的概念。为提升可靠性,引入概念保真度监督、先验一致性对齐、KL分布匹配与不确定性校准。在自建及LANL数据集上,该模块准确分类艾滋病毒亚型,并提供可直接与生物学先验验证的可解释证据。成本感知推荐层进一步将预测结果转化为兼顾准确性、校准性与临床效用的决策策略,减少不必要的复检,提升效率。大量实验表明,该系统在分类性能、概念预测保真度及成本效益权衡上均优于现有基线。本工作为基因组医学中的机器人与临床自动化建立了可靠基础。
原文摘要 · Abstract (English)
We propose an automated genomic interpretation module that transforms raw DNA sequences into actionable, interpretable decisions suitable for integration into medical automation and robotic systems. Our framework combines Chaos Game Representation (CGR) with a Concept Bottleneck Model (CBM), enforcing predictions to flow through biologically meaningful concepts such as GC content, CpG density, and k mer motifs. To enhance reliability, we incorporate concept fidelity supervision, prior consistency alignment, KL distribution matching, and uncertainty calibration. Beyond accurate classification of HIV subtypes across both in-house and LANL datasets, our module delivers interpretable evidence that can be directly validated against biological priors. A cost aware recommendation layer further translates predictive outputs into decision policies that balance accuracy, calibration, and clinical utility, reducing unnecessary retests and improving efficiency. Extensive experiments demonstrate that the proposed system achieves state of the art classification performance, superior concept prediction fidelity, and more favorable cost benefit trade-offs compared to existing baselines. By bridging the gap between interpretable genomic modeling and automated decision-making, this work establishes a reliable foundation for robotic and clinical automation in genomic medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。