arXiv:2607.09880cs.CLcs.AI2026-07被引 2

构建首个针对不规则临床时间序列的问答评测基准,解决模型难定位稀疏时间证据的问题。

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

论文配图:CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
图 1 · 摘自论文原文
  • 基于去标识化ICU数据,四阶段流程构建不规则时间序列问答集
  • 包含6600个问题,覆盖11个临床变量,验证模型对时间证据的使用能力
  • 揭示当前通用模型在稀疏时间推理上表现不佳,适合医疗AI研究者参考

临床时间序列在患者监测、风险评估和临床决策支持中至关重要,但常呈现稀疏、不规则采样和异步特点,使模型难以定位问答所需的时序证据。现有基准多聚焦于规则采样时间序列或静态医疗数据问答,很少评估模型是否能准确依据不规则时序观测得出答案。为填补这一空白,我们提出CLIR-Bench,一个基于去标识化ICU记录构建的不规则临床时间序列问答基准,采用系统性四阶段流程。该基准包含6,600个问答实例,涵盖11个临床变量,分为四个能力维度与11项任务。每个问题均关联明确的时间证据及任务特定的答案推导规则,可评估答案准确性与证据使用情况。实验表明,现有通用模型在检索和推理稀疏临床证据方面表现不佳,凸显发展更强不规则时间序列推理方法的必要性。代码与数据已开源于https://huggingface.co/datasets/winall/CLIR-Bench。

原文摘要 · Abstract (English)

Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA). Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. CLIR-Bench contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks. Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use. Experiments show that existing generalist models struggle to retrieve and reason over sparse clinical evidence, highlighting the need for stronger irregular time-series reasoning methods. Our code and data are available at https://huggingface.co/datasets/winall/CLIR-Bench.

临床问答时间序列医疗AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。