让医学数据的特征更丰富稳定,提升模型泛化能力
Dense Feature Learning via Linear Structure Preservation in Medical Data
- 通过保持嵌入矩阵的线性结构,显式优化特征分布
- 在纵向病历、临床文本等数据上显著提升特征稳定性与下游性能
- 无需标签或重建任务,适合追求可解释性的医疗AI研究
针对医学数据深度学习模型通常只关注特定任务目标,导致特征向量坍缩至少数判别方向的问题,本文提出密集特征学习(Dense Feature Learning),一种以表示为中心的框架,直接作用于嵌入矩阵,通过线性代数性质(谱平衡、子空间一致性、特征正交性)显式保留临床数据的丰富结构。该方法不依赖标签或生成重建,可生成更高有效秩、更好条件性和跨时间更稳定的表示。在纵向电子健康记录、临床文本及多模态患者表征上的实证表明,其下游线性任务表现、鲁棒性与子空间对齐均优于监督与自监督基线。结果提示:学习覆盖临床变异可能与预测临床结局同等重要,应将表示几何作为医疗AI的核心优化目标。
原文摘要 · Abstract (English)
Deep learning models for medical data are typically trained using task specific objectives that encourage representations to collapse onto a small number of discriminative directions. While effective for individual prediction problems, this paradigm underutilizes the rich structure of clinical data and limits the transferability, stability, and interpretability of learned features. In this work, we propose dense feature learning, a representation centric framework that explicitly shapes the linear structure of medical embeddings. Our approach operates directly on embedding matrices, encouraging spectral balance, subspace consistency, and feature orthogonality through objectives defined entirely in terms of linear algebraic properties. Without relying on labels or generative reconstruction, dense feature learning produces representations with higher effective rank, improved conditioning, and greater stability across time. Empirical evaluations across longitudinal EHR data, clinical text, and multimodal patient representations demonstrate consistent improvements in downstream linear performance, robustness, and subspace alignment compared to supervised and self supervised baselines. These results suggest that learning to span clinical variation may be as important as learning to predict clinical outcomes, and position representation geometry as a first class objective in medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。