arXiv:2609.06080cs.LGcs.AI2026-09

构建可复现的多模态健康数据评测基准,揭示不同指标对临床问题的预测价值。

PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

论文配图:PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
图 1 · 摘自论文原文
  • 以13,000人队列为基础,设计90个临床相关任务,统一评估标准
  • 发现测量值的预测能力因任务和表征方式而异,部分任务性能优于基线
  • 支持新模型、新问题快速接入,实现可追溯、可比较的持续评估

深度表型队列整合了从秒到年的临床、影像、分子及可穿戴设备数据。这种多样性可揭示哪些测量值能回答哪些健康问题,但异质分析难以直接比较。本文提出PhenoBench,一个基于人类表型项目(Human Phenotype Project)的可执行基准,该队列中超过13,000名参与者已完成初访。每个问题固定目标、适用人群、时间窗口与可用信息;其评估契约明确划分方式、指标、基线与声明边界。基准涵盖15个领域、26种输入模态的90个临床驱动任务。结果显示,测量值的预测价值具有任务与表征依赖性,部分任务在留出集上表现优于匹配基线,也有近零或负向变化。我们使用PhenoBench评估了160次匹配回归比较中的新兴表格基础模型,覆盖52个任务。这些模型总体优于专用模型,但平均仅比岭回归提升0.004 $R^2$(95% CI: 0.002–0.006)。随后,我们用相同队列数据和评估契约测试14个语言模型,覆盖40个任务,包括表型恢复、分类、随访预测与参与者排序。未经队列特定微调的语言模型在部分任务上仍具信息量,但存在任务特异性能力差距,共享规模缺陷,且极少超越同领域微调模型。PhenoBench将多模态纵向队列转化为版本化、可审计的评估系统,允许新增问题、测量与模型,无需重定义已有对比。

原文摘要 · Abstract (English)

Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent predictive value, including positive, near-zero, and negative changes in held-out performance relative to matched baselines. We used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate but, averaged across the three pretrained models within each cell, improved on ridge by a median of only 0.004 $R^2$ (95% CI, 0.002--0.006). We then used the same cohort data and evaluation contracts to evaluate 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.

多模态健康预测评测基准表型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。