arXiv:2507.01032cs.LGcs.AI2025-07被引 3

根据不确定性动态选择测序数据,省钱又准

An Uncertainty-Aware Dynamic Decision Framework for Progressive Multi-Omics Integration in Classification Tasks

  • 用信念度和不确定性量化分类结果,更可信
  • 3个数据集仅靠单一组学就超50%准确率
  • 适合资源有限但要精准诊断的医学研究

高通量多组学技术在揭示疾病机制和早期诊断中至关重要,但其高昂成本带来经济负担。为减少不必要的检测,本文提出一种考虑不确定性的动态决策框架,用于多组学分类。在单组学层面,通过改进神经网络激活函数生成狄利克雷分布参数,结合主观逻辑量化分类结果的信念质量与不确定性,其中信念反映特定组学对疾病类别的支持,不确定性捕捉数据质量与模型区分能力的局限。在多组学层面,采用基于达姆斯-谢弗理论的融合策略,整合异构模态以提升诊断准确性与鲁棒性。引入动态决策机制:逐个加入患者组学数据,直到模型置信度超过预设阈值或全部数据用完,部分病例可提前终止检测。在ROSMAP、LGG、BRCA和KIPAN四个基准数据集上评估,三个数据集中超过50%的样本仅使用单一组学即可实现准确分类,显著降低冗余检测,同时保持与全组学模型相当的性能并保留关键生物学信息。

原文摘要 · Abstract (English)

Background and Objective: High-throughput multi-omics technologies have proven invaluable for elucidating disease mechanisms and enabling early diagnosis. However, the high cost of multi-omics profiling imposes a significant economic burden, with over reliance on full omics data potentially leading to unnecessary resource consumption. To address these issues, we propose an uncertainty-aware, multi-view dynamic decision framework for omics data classification that aims to achieve high diagnostic accuracy while minimizing testing costs. Methodology: At the single-omics level, we refine the activation functions of neural networks to generate Dirichlet distribution parameters, utilizing subjective logic to quantify both the belief masses and uncertainty mass of classification results. Belief mass reflects the support of a specific omics modality for a disease class, while the uncertainty parameter captures limitations in data quality and model discriminability, providing a more trustworthy basis for decision-making. At the multi omics level, we employ a fusion strategy based on Dempster-Shafer theory to integrate heterogeneous modalities, leveraging their complementarity to boost diagnostic accuracy and robustness. A dynamic decision mechanism is then applied that omics data are incrementally introduced for each patient until either all data sources are utilized or the model confidence exceeds a predefined threshold, potentially before all data sources are utilized. Results and Conclusion: We evaluate our approach on four benchmark multi-omics datasets, ROSMAP, LGG, BRCA, and KIPAN. In three datasets, over 50% of cases achieved accurate classification using a single omics modality, effectively reducing redundant testing. Meanwhile, our method maintains diagnostic performance comparable to full-omics models and preserves essential biological insights.

多组学动态决策不确定性建模医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。