arXiv:2604.02834cs.AI2026-04

构建可编程的合成健康数据集,用于评测长期医疗智能体在多源数据下的推理能力。

ESL-Bench: An Event-Driven Synthetic Longitudinal Benchmark for Health Agents

  • 基于事件驱动框架生成100名用户、1-5年跨度的多模态健康轨迹。
  • 真实答案可程序化计算,支持对趋势、异常、解释等5类任务的精准评估。
  • 揭示数据库型智能体在复杂推理任务上显著优于记忆增强模型。

纵向健康智能体需整合连续设备流、稀疏临床检查和偶发生活事件等多种来源的轨迹数据,但其评估困难:真实数据难以大规模释放,且带时间锚点的归因问题缺乏结构化真值。本文提出ESL-Bench,一个事件驱动的合成框架与基准,包含100个合成用户,每个拥有1至5年的轨迹,涵盖健康档案、多阶段叙事计划、每日设备测量、周期性检查记录及带有显式指标影响参数的事件日志。每个指标遵循基础随机过程,由离散事件触发,采用S形起始、指数衰减核,在饱和与投影约束下演化;混合流水线将稀疏语义内容交由基于LLM的规划,密集指标动态则由满足生理边界条件的算法模拟。每位用户对应100个评估查询,覆盖查找、趋势、比较、异常、解释五个维度,分易、中、难三档,所有真值答案均可从事件-指标关系中程序化推导。对13种方法(包括带工具的LLM、原生数据库代理、记忆增强RAG)的评估显示,数据库代理(48-58%)显著优于记忆增强基线(30-38%),差距集中于需多跳推理与证据归因的比较与解释任务。

原文摘要 · Abstract (English)

Longitudinal health agents must reason across multi-source trajectories that combine continuous device streams, sparse clinical exams, and episodic life events - yet evaluating them is hard: real-world data cannot be released at scale, and temporally grounded attribution questions seldom admit definitive answers without structured ground truth. We present ESL-Bench, an event-driven synthesis framework and benchmark providing 100 synthetic users, each with a 1-5 year trajectory comprising a health profile, a multi-phase narrative plan, daily device measurements, periodic exam records, and an event log with explicit per-indicator impact parameters. Each indicator follows a baseline stochastic process driven by discrete events with sigmoid-onset, exponential-decay kernels under saturation and projection constraints; a hybrid pipeline delegates sparse semantic artifacts to LLM-based planning and dense indicator dynamics to algorithmic simulation with hard physiological bounds. Users are each paired with 100 evaluation queries across five dimensions - Lookup, Trend, Comparison, Anomaly, Explanation - stratified into Easy, Medium, and Hard tiers, with all ground-truth answers programmatically computable from the recorded event-indicator relationships. Evaluating 13 methods spanning LLMs with tools, DB-native agents, and memory-augmented RAG, we find that DB agents (48-58%) substantially outperform memory RAG baselines (30-38%), with the gap concentrated on Comparison and Explanation queries where multi-hop reasoning and evidence attribution are required.

健康代理合成数据多模态推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。