用知识图谱动态生成临床指南测试题,让大模型评估更全面可靠。
From Guidelines to Guarantees: A Graph-Based Evaluation Harness for Domain-Specific Evaluation of LLMs
- 将临床指南转为可查询的知识图谱,通过图遍历生成测试题。
- 模型在症状识别上准确率高,但治疗方案和决策能力明显不足。
- 适合医疗AI评估,支持指南更新后自动重生成测试数据。
领域特定语言模型的严格评估需要全面、抗污染且可维护的基准。静态人工标注数据集无法满足这些要求。我们提出一种基于图的评估框架,将结构化临床指南转化为可查询的知识图谱,并通过图遍历动态生成评估问题。该框架提供三项保障:(1)覆盖指南中所有关系;(2)通过组合变化抵抗表面形式污染;(3)从专家编写的图结构继承有效性。应用于世界卫生组织小儿传染病管理指南(WHO IMCI),该框架生成涵盖症状识别、治疗、严重程度分类和随访护理的临床真实多选题。在五种语言模型上的评估揭示系统性能力差距:模型在症状识别上表现良好,但在治疗方案和临床管理决策上准确率较低。该框架支持随指南演进持续再生评估数据,适用于具有结构化决策逻辑的其他领域,为可扩展的评估基础设施提供支持。
原文摘要 · Abstract (English)
Rigorous evaluation of domain-specific language models requires benchmarks that are comprehensive, contamination-resistant, and maintainable. Static, manually curated datasets do not satisfy these properties. We present a graph-based evaluation harness that transforms structured clinical guidelines into a queryable knowledge graph and dynamically instantiates evaluation queries via graph traversal. The framework provides three guarantees: (1) complete coverage of guideline relationships; (2) surface-form contamination resistance through combinatorial variation; and (3) validity inherited from expert-authored graph structure. Applied to the WHO IMCI guidelines, the harness generates clinically grounded multiple-choice questions spanning symptom recognition, treatment, severity classification, and follow-up care. Evaluation across five language models reveals systematic capability gaps. Models perform well on symptom recognition but show lower accuracy on treatment protocols and clinical management decisions. The framework supports continuous regeneration of evaluation data as guidelines evolve and generalizes to domains with structured decision logic. This provides a scalable foundation for evaluation infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。