针对医疗大模型评估难题,提出分层拆解评估框架提升与医生判断的一致性。
Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation
- 将复杂医疗评估拆分为患者问题相关性、医学知识正确性等子任务
- 通过专家模型与偏好数据训练,使评估结果与医生判断更一致
- 适合医疗AI研发者和临床评估人员参考使用
在大型语言模型(LLM)应用于医疗的快速发展中,确保其在临床场景中的可靠性与准确性至关重要。现有基准多聚焦于固定格式的任务(如选择题问答),难以反映真实临床诊断的复杂性。传统评估指标和基于LLM的评估器常因对齐不足,提供过于简化的判断,无法充分反映人类医生的评价。为此,我们提出一种面向细粒度对齐的分层分解评估框架HDCEval。该框架基于与专业医生协作制定的细粒度评估指南,涵盖患者问题相关性、医学知识正确性和表达质量。通过将复杂评估任务分解为专业化子任务,并利用属性驱动的词元优化(ADTO)在精心构建的偏好数据集上训练专家模型,实现各维度的精准评估,显著提升了评估结果与人类评判的一致性。
原文摘要 · Abstract (English)
In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of real-world clinical diagnostics. Moreover, traditional evaluation metrics and LLM-based evaluators struggle with misalignment, often providing oversimplified assessments that do not adequately reflect human judgment. To address these challenges, we introduce HDCEval, a Hierarchical Divide-and-Conquer Evaluation framework tailored for fine-grained alignment in medical evaluation. HDCEval is built on a set of fine-grained medical evaluation guidelines developed in collaboration with professional doctors, encompassing Patient Question Relevance, Medical Knowledge Correctness, and Expression. The framework decomposes complex evaluation tasks into specialized subtasks, each evaluated by expert models trained through Attribute-Driven Token Optimization (ADTO) on a meticulously curated preference dataset. This hierarchical approach ensures that each aspect of the evaluation is handled with expert precision, leading to a significant improvement in alignment with human evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。