用常规化验数据预测犬类癌症风险,模型有统计信号但临床不可靠。
Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset
- 用逻辑回归+加权+递归特征选择,在犬类化验数据中建模癌症风险。
- 模型AUROC达0.815,但准确率仅F1=0.25,阳性预测值仅0.15。
- 主要依赖年龄和炎症等非特异性指标,难区分癌症与衰老或炎症。
犬类早期癌症筛查工具的开发在兽医学中面临重大挑战。常规实验室数据为低成本筛查提供了潜在可能,但受限于单个生物标志物的非特异性以及筛查人群中的严重类别不平衡。本研究在金毛寻回猎犬终身研究(GRLS)队列中评估了机器学习方法在现实约束下的可行性,包括多种癌症类型的合并及诊断后样本的纳入。系统比较了126种分析流程,涵盖不同机器学习模型、特征选择方法与数据平衡技术,按患者级别划分数据以防泄露。最优模型为带类别权重的逻辑回归与递归特征消除组合,表现出中等排序能力(AUROC = 0.815;95% CI: 0.793–0.836),但临床分类性能差(F1-score = 0.25,阳性预测值 = 0.15)。尽管获得高阴性预测值(0.98),但召回率不足(0.79),无法作为可靠的排除测试。基于SHAP的可解释性分析显示,预测主要由年龄及炎症、贫血等非特异性指标驱动。结论表明,尽管常规化验数据中存在可统计检测的癌症信号,但其强度过弱且易受正常老化或其他炎症干扰,难以实现临床可靠区分。该研究确立了仅依赖此类数据的性能上限,并强调计算兽医肿瘤学的实质性进展需融合多模态数据源。
原文摘要 · Abstract (English)
The development of accessible screening tools for early cancer detection in dogs represents a significant challenge in veterinary medicine. Routine laboratory data offer a promising, low-cost source for such tools, but their utility is hampered by the non-specificity of individual biomarkers and the severe class imbalance inherent in screening populations. This study assesses the feasibility of cancer risk classification using the Golden Retriever Lifetime Study (GRLS) cohort under real-world constraints, including the grouping of diverse cancer types and the inclusion of post-diagnosis samples. A comprehensive benchmark evaluation was conducted, systematically comparing 126 analytical pipelines that comprised various machine learning models, feature selection methods, and data balancing techniques. Data were partitioned at the patient level to prevent leakage. The optimal model, a Logistic Regression classifier with class weighting and recursive feature elimination, demonstrated moderate ranking ability (AUROC = 0.815; 95% CI: 0.793-0.836) but poor clinical classification performance (F1-score = 0.25, Positive Predictive Value = 0.15). While a high Negative Predictive Value (0.98) was achieved, insufficient recall (0.79) precludes its use as a reliable rule-out test. Interpretability analysis with SHapley Additive exPlanations (SHAP) revealed that predictions were driven by non-specific features like age and markers of inflammation and anemia. It is concluded that while a statistically detectable cancer signal exists in routine lab data, it is too weak and confounded for clinically reliable discrimination from normal aging or other inflammatory conditions. This work establishes a critical performance ceiling for this data modality in isolation and underscores that meaningful progress in computational veterinary oncology will require integration of multi-modal data sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。