让大模型动态调用工具,精准回答可穿戴设备的健康问题
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning

- 根据问题自动选择分析工具和模型组合,实现动态推理
- 在四个公开数据集上比基线高24%准确率,专家评测效果显著提升
- 适合医疗健康领域研究者与可穿戴设备开发者使用
大语言模型在医学问答中表现优异,甚至超过普通医生。但针对可穿戴设备产生的连续、高维、长期的健康数据,传统方法仍面临挑战,因这些数据难以与文本主导预训练的LLM对齐。不同传感器类型和用户意图无法通过固定推理流程或单一基础模型有效处理。为此,我们提出WEQA,一种查询自适应的智能体框架,将LLM推理与专用可穿戴数据分析工具及预训练模型统一。由LLM控制器生成执行计划,动态调度查询至合适的分析与模型组合,并利用外部知识进行响应验证。我们还构建了涵盖四个公开可穿戴数据集的基准,覆盖三个健康领域中的分析与预测任务。实验表明,该框架比LLM与代理基线高出24%准确率;12名医学专家与8名用户的盲评显示其在实用性和临床合理性上均有显著提升。
原文摘要 · Abstract (English)
Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However, answering questions about wearable health data remains challenging and understudied, as these ubiquitous sensors produce continuous, high-dimensional, and longitudinal data, which is non-trivial to align with text-centric distributions in LLM pretraining. The diversity of sensor modalities and user intents cannot be effectively handled by a fixed reasoning workflow or a single pretrained foundation model. To address these challenges, we propose WEQA, a query-adaptive agent framework that unifies LLM reasoning with specialized wearable analytical and modeling tools. An LLM controller is employed to synthesize execution plans and dynamically route each query to the appropriate combination of sensor analysis and pretrained models, and perform grounded response auditing with external knowledge. We also curate a benchmark spanning four open wearable datasets comprising analytic and predictive tasks in three different health domains. Experiments show that our framework is 24% more accurate than LLM and agentic baselines, and a blinded study with 12 medical experts and 8 users shows substantial gains in usefulness and clinical soundness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。