arXiv:2607.25947cs.AIcs.CL2026-07被引 1

用少到16个时间序列令牌,实现高效精准的临床时间序列问答。

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

论文配图:A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
图 1 · 摘自论文原文
  • 设计多尺度编码器捕捉临床数据的稀疏性与异步性
  • 仅用16个令牌完成推理,平均响应速度0.15秒
  • 适合医疗智能问答系统快速部署与低资源应用

面向不规则临床时间序列(ICTS)的问答在多种医疗应用中至关重要。尽管近期多模态时间序列大语言模型在通用时间序列问答上表现良好,但对临床观察中普遍存在的稀疏性、异步性和非均匀采样模式仍缺乏有效建模。为此,我们提出ClinPRISM——一种低成本、高效的多模态大模型推理框架。首先,设计了考虑不规则性的多尺度编码器,在不同时间粒度下捕获稀疏临床证据;其次,提出时序证据提炼器,将多尺度表示压缩为少量适配大模型的令牌;此外,引入渐进对齐策略,逐步将不规则轨迹映射至大模型的文本嵌入空间。为支持训练,构建了3万条带多尺度描述的临床时间序列数据,以及4.1万条覆盖11项任务的指令微调实例。基于40亿参数的大模型骨干,ClinPRISM在留出测试集上达到当前最优性能,仅使用16个时间序列令牌,且单问题平均推理延迟为0.15秒。

原文摘要 · Abstract (English)

Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.

医疗AI多模态大模型推理时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。