构建真实可穿戴数据上的健康推理基准,评测AI理解长期生理数据的能力。
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

- 基于200人500天的可穿戴数据构建4084道多选题
- 模型在真实数据上准确率仅19.6%至72.9%,远低于人类水平
- 涵盖16类推理任务,适合评估医疗大模型的逻辑与整合能力
可穿戴传感技术实现对生理与行为信号的持续监测,但现有基准很少评估AI能否对真实用户的长期可穿戴记录进行推理。我们提出WearableQA,一个包含4,084道10选项多选题的基准,数据来自200名真实用户,每人最多500天的每日测量,涵盖可穿戴时间序列、血液生物标志物和人口统计信息。该基准保留了真实可穿戴数据的分布特征,包括设备噪声和个体差异。为评估不同推理能力,我们设计了16种问题类型,沿两个互补维度划分:数据与健康推理(区分对纵向数据的计算与生理解释);单信号与跨信号推理(区分单一信号分析与多信号融合)。通过结合文献依据的生理发现与统计验证的人群模式,采用双重锚定框架实现大规模可靠问题构建,捕捉真实可穿戴数据中的有意义关系。对14个专有及开源大模型的评估表明,WearableQA能有效区分模型能力,性能范围为19.6%至72.9%(随机基线为10%)。此外,该任务仍远未解决:多数模型准确率低于60%。总体而言,WearableQA为评估大模型在真实可穿戴数据上的推理能力提供了现实且具有诊断性的基准。
原文摘要 · Abstract (English)
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。