arXiv:2603.01557cs.AI2026-03被引 1

评估大模型对临床时间序列的总结是否准确捕捉关键异常事件。

Benchmarking LLM Summaries of Multimodal Clinical Time Series for Remote Monitoring

  • 基于规则定义临床事件,用结构化事实比对生成摘要
  • 视觉化方法在异常事件召回率达45.7%,持续时间召回100%
  • 提醒研究者关注事件级准确率,而非仅看语义相似度

大型语言模型(LLMs)可生成远程治疗监测时间序列的流畅临床摘要。然而,这些叙述是否真实反映临床重要事件(如持续异常)仍不明确。现有评估指标主要关注语义相似性和语言质量,未衡量事件层面的正确性。为此,我们引入基于事件的评估框架,使用Technology-Integrated Health Management (TIHM)-1.5痴呆监测数据集。通过规则设定异常阈值和时间持续性标准,提取每日临床事件。将模型生成的摘要与这些结构化事实对齐,评估异常事件召回率、持续时间召回率、测量覆盖率及幻觉事件提及。我们对比三种方法:零样本提示、统计提示和基于视觉化的流水线。结果表明,传统指标与临床事件准确性严重脱节:高语义相似度模型异常事件召回率接近零。而视觉化方法表现最优,异常事件召回率达45.7%,持续时间召回率为100%。研究强调了事件感知评估对可靠临床时间序列摘要的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) can generate fluent clinical summaries of remote therapeutic monitoring time series. However, it remains unclear whether these narratives faithfully capture clinically significant events, such as sustained abnormalities. Existing evaluation metrics primarily focus on semantic similarity and linguistic quality, leaving event-level correctness largely unmeasured. To address this gap, we introduce an event-based evaluation framework for multimodal time-series summarization using the Technology-Integrated Health Management (TIHM)-1.5 dementia monitoring dataset. Clinically grounded daily events are derived through rule-based abnormal thresholds and temporal persistence criteria. Model-generated summaries are then aligned with these structured facts. Our evaluation protocol measures abnormality recall, duration recall, measurement coverage, and hallucinated event mentions. We benchmark three approaches: zero-shot prompting, statistical prompting, and a vision-based pipeline that uses rendered time-series visualizations. The results reveal a striking decoupling between conventional metrics and clinical event fidelity. Models that achieve high semantic similarity scores often exhibit near-zero abnormality recall. In contrast, the vision-based approach demonstrates the strongest event alignment, achieving 45.7% abnormality recall and 100% duration recall. These findings underscore the importance of event-aware evaluation to ensure reliable clinical time-series summarization.

临床摘要时间序列大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。