复杂算法无法弥补基础设计缺陷,先验方法一致性才是关键。
Methodological Precedence in Health Tech: Why ML/Big Data Analysis Must Follow Basic Epidemiological Consistency. A Case Study
- 用基础统计方法检验疫苗研究数据,发现根本性设计问题
- 暴露组癌症率上升与总体率下降矛盾,揭示数据偏差
- 适合关注研究可信度与方法论的临床与公共卫生学者
尽管机器学习与大数据分析在健康研究中带来诊断与风险预测的突破,但其严谨性依赖于数据质量与统计设计的有效性。本研究通过一个新冠疫苗结局与严重不良事件(如癌症)的队列研究案例,强调先进分析必须建立在基本流行病学一致性的基础上。采用标准描述性统计与国家流行病学基准,我们发现暴露组癌症发病率上升,而总人群粗发病率却低于国家标准,两者存在不可调和的矛盾,证明癌症风险增加的结论无效。该现象源于队列构建中的未校正选择偏倚,是数学伪象。结果表明,任何高级建模前,研究必须通过基本方法学一致性检验,否则结论不可信。
原文摘要 · Abstract (English)
The integration of advanced analytical tools, including Machine Learning (ML) and massive data processing, has revolutionized health research, promising unprecedented accuracy in diagnosis and risk prediction. However, the rigor of these complex methods is fundamentally dependent on the quality and integrity of the underlying datasets and the validity of their statistical design. We propose an emblematic case where advanced analysis (ML/Big Data) must necessarily be subsequent to the verification of basic methodological coherence and adherence to established medical protocols, such as the STROBE Statement. This study highlights a crucial cautionary principle: sophisticated analyses amplify, rather than correct, severe methodological flaws rooted in basic design choices, leading to misleading or contradictory findings. By applying simple, standard descriptive statistical methods and established national epidemiological benchmarks to a recently published cohort study on COVID-19 vaccine outcomes and severe adverse events, like cancer, we expose multiple, statistically irreconcilable paradoxes. These paradoxes, specifically the contradictory finding of an increased cancer incidence within an exposure subgroup, concurrent with a suppressed overall Crude Incidence Rate compared to national standards, definitively invalidate the reported risk of increased cancer in the total population. We demonstrate that the observed effects are mathematical artifacts stemming from an uncorrected selection bias in the cohort construction. This analysis serves as a robust reminder that even the most complex health studies must first pass the test of basic epidemiological consistency before any conclusion drawn from subsequent advanced statistical modeling can be considered valid or publishable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。