arXiv:2603.06638cs.LGcs.AI2026-03被引 7

构建首个覆盖12大领域、20种信号的医疗时序推理基准

HEARTS: Benchmarking LLM Reasoning on Health Time Series

  • 整合16个真实数据集,定义感知/推断/生成/推理四类能力
  • 16个主流大模型在超2万样本上测试,表现普遍弱于专用模型
  • 揭示模型依赖简单规则、难处理多步时序推理的共性缺陷

大型语言模型(LLMs)的兴起使时间序列分析从狭义分析转向通用推理。然而现有基准仅涵盖少量医疗时序模态与任务,难以反映真实生理建模中多样的领域和复杂的时序依赖。为此,我们提出HEARTS(Health Reasoning over Time Series),一个统一的基准,用于评估LLMs在通用医疗时序上的层次化推理能力。HEARTS整合了16个真实世界数据集,覆盖12个健康领域和20种信号模态,并定义了包含110项任务的综合性分类体系,分为感知、推断、生成和推理四大核心能力。在超过2万测试样本上对16个先进LLMs进行评估发现:首先,LLMs显著低于专用模型,其性能与通用推理得分关联微弱;其次,模型常依赖简单启发式规则,难以处理多步时序推理;最后,随着时序复杂度增加,性能下降,同族模型表现出相似失败模式,表明单纯扩大规模不足以提升能力。通过量化这些差距,HEARTS为开发下一代能处理多样化健康信号的LLM智能体提供了标准化测试平台与动态基准。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing benchmarks cover only a small set of health time series modalities and tasks, failing to reflect the diverse domains and extensive temporal dependencies inherent in real-world physiological modeling. To bridge these gaps, we introduce HEARTS (Health Reasoning over Time Series), a unified benchmark for evaluating hierarchical reasoning capabilities of LLMs over general health time series. HEARTS integrates 16 real-world datasets across 12 health domains and 20 signal modalities, and defines a comprehensive taxonomy of 110 tasks grouped into four core capabilities: Perception, Inference, Generation, and Deduction. Evaluating 16 state-of-the-art LLMs on more than 20K test samples reveals intriguing findings. First, LLMs substantially underperform specialized models, and their performance is only weakly related to general reasoning scores. Moreover, LLMs often rely on simple heuristics and struggle with multi-step temporal reasoning. Finally, performance declines with increasing temporal complexity, with similar failure modes within model families, indicating that scaling alone is insufficient. By making these gaps measurable, HEARTS provides a standardized testbed and living benchmark for developing next-generation LLM agents capable of reasoning over diverse health signals.

医疗时序大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。