将无序生物医学表格数据转为可视空间图谱,提升模型预测精度。
Vision-based Deep Learning Analysis of Unordered Biomedical Tabular Datasets via Optimal Spatial Cartography
- 通过可微渲染机制动态学习特征空间布局,无需预设分组或先验知识。
- 在液体活检数据中提升癌症亚型分类准确率最高达18%,语音数据提升8%。
- 适合需要可解释性与高维数据建模的临床研究者和生物信息学专家。
表格数据在生物医学研究中至关重要,涵盖液体活检、批量及单细胞转录组、电子健康记录和表型分析等。然而,与图像或序列不同,表格数据缺乏内在空间结构:特征被视为无序维度,其关系需由模型隐式推断,限制了视觉架构对局部结构和高阶特征交互的利用。本文提出动态特征映射(Dynomap),一种端到端深度学习框架,可直接从数据中学习任务优化的特征拓扑。Dynomap通过全可微渲染机制联合优化特征位置与预测结果,无需启发式规则、预定义分组或外部先验。通过将高维表格向量转化为学习得到的特征图,使视觉模型能有效处理无序生物医学输入。在多个临床与生物数据集上,Dynomap持续优于传统机器学习、现代深度表格模型及现有向量到图像方法。在液体活检数据中,它将临床相关的基因标志物组织成连贯的空间模式,并将多分类癌症亚型预测准确率提升最高达18%;在帕金森病语音数据集中,成功聚类疾病相关声学描述符,准确率提升最高达8%。其他生物医学数据集也观察到类似性能提升与可解释的特征组织。这些结果确立了Dynomap作为连接表格与视觉深度学习的通用策略,以及在高维生物医学数据中发现结构化、临床相关模式的有效途径。
原文摘要 · Abstract (English)
Tabular data are central to biomedical research, from liquid biopsy and bulk and single-cell transcriptomics to electronic health records and phenotypic profiling. Unlike images or sequences, however, tabular datasets lack intrinsic spatial organization: features are treated as unordered dimensions, and their relationships must be inferred implicitly by the model. This limits the ability of vision architectures to exploit local structure and higher-order feature interactions in non-spatial biomedical data. Here we introduce Dynamic Feature Mapping (Dynomap), an end-to-end deep learning framework that learns a task-optimized spatial topology of features directly from data. Dynomap jointly optimizes feature placement and prediction through a fully differentiable rendering mechanism, without relying on heuristics, predefined groupings, or external priors. By transforming high-dimensional tabular vectors into learned feature maps, Dynomap enables vision-based models to operate effectively on unordered biomedical inputs. Across multiple clinical and biological datasets, Dynomap consistently outperformed classical machine learning, modern deep tabular models, and existing vector-to-image approaches. In liquid biopsy data, Dynomap organized clinically relevant gene signatures into coherent spatial patterns and improved multiclass cancer subtype prediction accuracy by up to 18%. In a Parkinson disease voice dataset, it clustered disease-associated acoustic descriptors and improved accuracy by up to 8%. Similar gains and interpretable feature organization were observed in additional biomedical datasets. These results establish Dynomap as a general strategy for bridging tabular and vision-based deep learning and for uncovering structured, clinically relevant patterns in high-dimensional biomedical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。