构建可扩展的医疗健康助手评估框架,提升AI诊断的准确性与临床一致性。
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

- 基于100多个可验证的医学评分规则,建立分层评估体系。
- 在4000条真实用户提问上迭代优化,专家对齐度显著提升。
- 适用于医疗AI模型训练、反馈和性能优化,支持多模型适配。
由大语言模型驱动的个人健康助手结合用户健康数据(如传感器指标),为缓解全球医疗资源不均提供了可行路径。然而,大规模临床应用仍受限于开放式评估瓶颈:医生标注虽可靠但成本高且难扩展;而以LLM为裁判的评估方法虽可扩展,却存在主观性、不一致性和临床偏差问题。本文提出RubricsTree,一个可扩展的评估框架,包含超过100个原子级、可临床验证的布尔评分规则,其结构源于4000条真实用户查询的迭代人机协同梳理过程,由经验丰富的医师领衔专家小组完成。通过上下文感知的自适应路由机制,仅激活相关且加权的评分子集,实现高效评估的同时保持专家对齐。系统性元评估表明,RubricsTree(i)在复杂开放式问题上显著优于强基准,专家对齐度更高;(ii)能可靠识别上下文退化的回答;(iii)作为结构化指令、文本反馈或训练奖励使用时,使Gemini、GPT和Qwen系列模型在HealthBench上的性能提升最高达约66%相对增益。RubricsTree为产品级医疗AI的持续优化提供了可扩展、可审计、可演进的评估基础设施。
原文摘要 · Abstract (English)
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access. However, large-scale clinical deployment remains constrained by an open-ended evaluation bottleneck: physician annotation is reliable but costly and unscalable, while LLM-as-a-judge evaluators are scalable but subjective, inconsistent, and sometimes clinically misaligned. We introduce RubricsTree, a scalable evaluation framework with an expert-aligned hierarchical taxonomy of over 100 atomic, clinically-verifiable Boolean rubrics, evolving from the insights of 4,000 real user queries through an iterative human-in-the-loop curation protocol with an expertise panel led by an experienced physician. A context-aware adaptive router activates only the relevant auto-weighted rubric subset per query, providing the throughput needed for scalable evaluation with expert-aligned quality. Through a systematic meta-evaluation, we show that RubricsTree (i) substantially exceeds a strong large-scale evaluation baseline in expert alignment on challenging open-ended queries; (ii) reliably penalizes contextually degraded responses; and (iii) when used as structured instructions, text feedback, or training rewards for performance optimization, yields up to ~66% relative gains on HealthBench for Gemini, GPT, and Qwen model families. RubricsTree thus provides a scalable, auditable, and evolving evaluation infrastructure required for the continuous optimization of product-level personal healthcare AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。