构建首个长周期跨维度健康推理基准,评测大模型真实健康分析能力。
LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning
- 设计多维度、长周期的问答基准,支持复杂推理任务
- 13个大模型评估揭示长期聚合与跨维度推理是主要瓶颈
- 提出工具增强的LifeAgent框架,提升复杂健康推理效果
个性化生活方式健康分析需要对异构生活信号进行长周期、多维度推理,移动传感和大语言模型(LLMs)的发展使这一目标日益可行。然而,由于缺乏系统性基准,当前LLMs在此场景下的能力仍不清晰。本文提出LifeAgentBench,一个大规模问答基准,涵盖22,573个问题,覆盖从基础检索到复杂推理的全谱任务。我们发布可扩展的基准构建流程与标准化评估协议,通过可执行查询和程序生成可验证答案,实现可靠评估。系统评估13个代表性LLMs,发现长期信息聚合与跨维度推理存在关键瓶颈。基于此,我们提出LifeAgent——一种工具增强的推理基线,通过分解复杂查询、多步证据检索并调用工具实现确定性聚合。该方法显著提升大模型在挑战性推理任务中的表现,优于主流基线,展现出在日常健康场景中的应用潜力。基准已公开可用。
原文摘要 · Abstract (English)
Personalized lifestyle health analysis requires long-horizon, multi-dimensional reasoning over heterogeneous lifestyle signals, and recent advances in mobile sensing and large language models (LLMs) make such support increasingly feasible. However, the capabilities of current LLMs in this setting remain insufficiently understood due to the lack of systematic benchmarks. In this paper, we introduce LifeAgentBench, a large-scale QA benchmark for long-horizon, cross-dimensional, and multi-user lifestyle health reasoning, containing 22,573 questions spanning from basic retrieval to complex reasoning. We release an extensible benchmark construction pipeline and a standardized evaluation protocol, deriving verifiable answers through executable queries and programs to support reliable assessment. We then systematically evaluate 13 representative LLMs on LifeAgentBench and identify key bottlenecks in long-horizon aggregation and cross-dimensional reasoning. Motivated by these findings, we propose LifeAgent, a tool-augmented reasoning baseline that decomposes complex queries, performs multi-step evidence retrieval, and invokes tools for deterministic aggregation. LifeAgent substantially enhances LLMs' capabilities on challenging reasoning tasks, achieving clear improvements over widely used baselines and showing potential for health reasoning in everyday scenarios. The benchmark is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。