arXiv:2608.00073cs.CVcs.LG2026-08

提出无监督时空分层方法,解决医学影像纵向数据分割偏差问题。

Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging

论文配图:Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging
图 1 · 摘自论文原文
  • 基于六维联合强度-时间特征空间的聚类与分层采样
  • 将最大跨子集强度偏差从34.1%降至2.1%以下
  • 适合纵向医学影像研究者用于提升模型可靠性

在纵向医学影像深度学习中,严谨的数据划分是基础但常被忽视的环节。随意打乱小规模临床队列会引入协变量偏移和时间采样不平衡,使下游模型面临分布外评估风险。本文提出可审计的三重数据集分析框架,系统刻画空间网格完整性、多参数强度指纹和纵向时间轨迹,量化真实临床队列中常见的重尾特征分布与不规则、间歇性采样间隔。在此基础上,建立无监督时空队列平衡标准操作流程:在标准化六维联合强度-时间特征空间中采用肘部优化的K均值聚类,并结合簇内比例分层采样。在包含149例增强T1加权脑MRI的纵向队列上,该方法将跨子集最大强度偏差从随机划分下的34.1%降至2.1%以下,同时使随访间隔紧密围绕人群均值。蒙特卡洛压力测试(10个随机种子、3种划分配置)表明,该结果保持高度稳定,显著优于随机划分的大幅波动。该协议为可重复、可推广的变长纵向临床影像队列构建提供新范式。

原文摘要 · Abstract (English)

Rigorous dataset partitioning is a foundational, yet frequently overlooked, prerequisite for reliable deep learning in longitudinal medical imaging. Naively shuffling small clinical cohorts routinely introduces covariate shifts and temporal sampling imbalances across training, validation, and test subsets, exposing downstream models to out-of-distribution evaluation. We address this vulnerability with an auditable Tripartite Dataset Analytics Framework that systematically characterizes spatial grid integrity, multi-parametric intensity fingerprints, and longitudinal temporal trajectories, quantifying the heavy-tailed feature dispersion and irregular, episodic sampling intervals typical of real-world clinical cohorts. Building on this characterization, we formalize an unsupervised spatio-temporal cohort-balancing standard operating procedure (SOP) that combines elbow-optimized K-means clustering over a standardized, six-dimensional joint intensity-temporal feature space with intra-cluster proportionate stratified sampling. On a longitudinal, contrast-enhanced $T1$-weighted brain MRI cohort (N=149), the protocol reduces the maximum cross-subset intensity bias from 34.1% under conventional random shuffling to under 2.1%, while aligning longitudinal follow-up intervals closely around the population mean. Monte Carlo stress testing across ten random seeds and three split configurations confirms that this alignment remains tightly bounded, in clear contrast to the substantial variability of random partitioning. The resulting protocol offers a reproducible, generalizable procedure for cohort engineering in variable-length longitudinal clinical imaging workflows.

医学影像纵向数据数据分割聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。