用精简布尔问答提升医疗大模型评估效率与一致性
A Scalable Framework for Evaluating Health Language Models
- 设计基于最小化精准布尔问题的评估框架,自动识别模型回答漏洞
- 相比传统评分法,专家与非专家评估一致性更高,耗时减少一半
- 适用于医疗健康领域,尤其适合非专业人员参与大规模模型评估
大型语言模型(LLMs)在分析复杂数据集方面展现出强大能力。近期研究显示,当提供包含生活方式、生物标志物和情境的患者特定健康信息时,LLMs 可生成有用且个性化的回应。随着基于 LLM 的健康应用日益普及,对生成文本进行严谨高效的单向评估至关重要,需覆盖准确性、个性化和安全性等多个维度。当前对开放式文本响应的评估严重依赖人工专家,存在成本高、耗时长、难以扩展的问题,尤其在需要领域知识的医疗领域更为突出。本文提出 Adaptive Precise Boolean rubrics:一种通过最少数量的针对性布尔问题识别模型回答缺失的评估框架。该方法借鉴通用评估场景中的思路,用少量复杂目标对比大量精细、可简单二值回答的目标。我们在代谢健康领域(涵盖糖尿病、心血管疾病和肥胖)验证了该方法,结果表明,相较于传统李克特量表,该框架在专家与非专家评估中均显著提高评分者间一致性,在自动化评估中也表现更优,且评估时间约为李克特法的一半。这一效率提升为医疗领域大规模、低成本的 LLM 评估提供了可行路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have emerged as powerful tools for analyzing complex datasets. Recent studies demonstrate their potential to generate useful, personalized responses when provided with patient-specific health information that encompasses lifestyle, biomarkers, and context. As LLM-driven health applications are increasingly adopted, rigorous and efficient one-sided evaluation methodologies are crucial to ensure response quality across multiple dimensions, including accuracy, personalization and safety. Current evaluation practices for open-ended text responses heavily rely on human experts. This approach introduces human factors and is often cost-prohibitive, labor-intensive, and hinders scalability, especially in complex domains like healthcare where response assessment necessitates domain expertise and considers multifaceted patient data. In this work, we introduce Adaptive Precise Boolean rubrics: an evaluation framework that streamlines human and automated evaluation of open-ended questions by identifying gaps in model responses using a minimal set of targeted rubrics questions. Our approach is based on recent work in more general evaluation settings that contrasts a smaller set of complex evaluation targets with a larger set of more precise, granular targets answerable with simple boolean responses. We validate this approach in metabolic health, a domain encompassing diabetes, cardiovascular disease, and obesity. Our results demonstrate that Adaptive Precise Boolean rubrics yield higher inter-rater agreement among expert and non-expert human evaluators, and in automated assessments, compared to traditional Likert scales, while requiring approximately half the evaluation time of Likert-based methods. This enhanced efficiency, particularly in automated evaluation and non-expert contributions, paves the way for more extensive and cost-effective evaluation of LLMs in health.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。