arXiv:2608.28605cs.AIcs.CV2026-08KDD

将视觉、语言与时间序列结合,提升医学时序分类准确性。

MedTVL: Harnessing Vision and Language for Medical Time Series Classification

论文配图:MedTVL: Harnessing Vision and Language for Medical Time Series Classification
图 1 · 摘自论文原文
  • 双路径架构融合时序数值与图像特征,分别捕捉细节与整体形态。
  • 引入文本引导和专家混合机制,实现个性化融合与诊断决策。
  • 适用于少样本和无标签场景,适合临床辅助诊断系统开发。

近年来,多模态学习在医学时序(MedTS)分类中的进展表明,整合互补模态有助于临床决策。然而,现有方法通常聚焦于双模态交互(如时序数据与文本),对时序、视觉与语言三者协同作用的研究仍不充分。受临床诊断中数值评估、视觉检查与临床语境结合的启发,我们提出 MedTVL,一种面向医学时序分类的文本引导双路径架构。该模型通过卷积路径捕捉原始数值序列中的精细时序动态,通过基于 Transformer 的视觉路径分析时序生成的图像所体现的全局形态结构。两者结合实现跨模态异构融合,提供全面诊断视角。为缓解潜在诊断模糊性,两条路径均受自适应医学文本语义引导。最后,采用混合专家机制动态分配实例至专用融合专家,捕获不同样本对时序与视觉路径输出的依赖差异。此外,MedTVL 支持多模态对比学习,以应对临床标签稀缺问题。在多个医学数据集和任务(包括监督、小样本与对比学习)上的广泛实验表明,MedTVL 具有优越性能与强泛化能力,展现出构建鲁棒临床决策支持系统的潜力。

原文摘要 · Abstract (English)

Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalities for clinical decision. However, existing methods typically focus on bi-modal interactions (e.g., time series and text), leaving the tri-modal synergy between time series, vision, and language largely unexplored. Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, we introduce MedTVL, a text-guided dual-pathway architecture tailored for MedTS classification. Specifically, it synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences and a transformer-based visual pathway for holistic morphological structures from time-series-derived images. Such combination of cross-modal and architectural heterogeneity provides a comprehensive diagnostic perspective. To further resolve potential diagnostic ambiguity, both pathways are guided by adaptive medical textual semantics. Finally, a Mixture-of-Experts mechanism dynamically routes each instance to specialized fusion experts, capturing instance-specific reliance on the temporal and visual pathway outputs. In addition, MedTVL supports multimodal contrastive learning to mitigate the clinical label scarcity challenge. Extensive experiments across multiple medical datasets and tasks, spanning supervised, few-shot, and contrastive learning settings, demonstrate the superiority and transferability of MedTVL, highlighting its potential for robust clinical decision support.

医学时序多模态融合双路径文本引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。