arXiv:2608.10522cs.CVcs.AI2026-08

让医学表格数据更聪明:用结构化语义提升模型表现

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

论文配图:Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
图 1 · 摘自论文原文
  • 设计双维度表格建模框架,显式捕捉特征重要性与连续离散关系
  • 在皮肤科和眼科数据集上达到新SOTA,跨域泛化能力强
  • 适合医疗表结构数据建模、临床决策支持系统研发者

尽管视觉-语言模型在医学表示学习中占据主导地位,但非结构化文本缺乏结构化临床表格中固有的密集定量诊断表型。现有多模态预训练方法因语义无关的设计,将表格输入视为扁平向量,并采用不稳定的连续回归目标,未能充分挖掘其潜力。为此,我们提出一种新的语义感知框架,显式建模表格数据的内在二维结构。首先,针对不同特征间诊断重要性的层次关系,引入重要性感知自适应掩码,构建无需标签的课程学习策略,优先关注关键特征;其次,针对特征内部连续性与离散性的双重特性,提出软标签离散化模块,以稳定的分布匹配替代不稳定的数值回归,从而数学上保持序数关系。在大规模皮肤病学(SLICE-3D、HOP)和眼科学(EyePACS)数据集上的大量实验表明,该方法达到新的最先进水平(SOTA),展现出优异的鲁棒性和跨领域泛化能力。

原文摘要 · Abstract (English)

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing multimodal pre-training methods underutilize this potential due to semantic-agnostic designs that treat tabular inputs as flat vectors and employ unstable continuous regression objectives. To overcome this, we propose a novel semantic-aware framework explicitly modeling the intrinsic two-dimensional structure of tabular data. First, addressing the inter-feature hierarchy of varying diagnostic importance, we introduce Importance-Aware Adaptive Masking to construct a label-free curriculum prioritizing salient features. Second, addressing the intra-feature continuity-discreteness duality, we propose a Soft-Label Discretized Module that replaces unstable numerical regression with stable distribution matching, thereby mathematically preserving ordinal relationships. Extensive experiments across large-scale dermatology (SLICE-3D, HOP) and ophthalmology (EyePACS) datasets establish a new state-of-the-art (SOTA), demonstrating exceptional robustness and cross-domain generalizability.

医学表数据多模态预训练结构化学习临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。