构建可追溯的图证据基准,提升罕见病诊断模型透明度
GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis

- 设计含图定义混淆因子的可审计基准,追踪模型决策依据
- 模型在237例测试中平均排序准确率0.640~0.740,目标证据覆盖差异达0.561
- 适合评估诊断模型是否真正理解临床证据,尤其关注误判根源
表型驱动的诊断基准通常仅报告参考疾病的排名,却很少揭示哪些合理替代疾病排在其前,或模型决策前依赖何种证据。我们提出GraphRareBench,一个保留溯源信息的基准,包含2,365个基于本体的病例和18,093个目标-混淆对。每例包含粗粒度HPO查询、固定候选池、图定义的硬混淆因子及来源关联的证据记录。在237例基因组组件不交集的测试集上,使用共享21特征接口的监督排序器取得MRR 0.640~0.740,案例平均目标-混淆准确率0.898~0.916。采用Agents-A1和DeepSeek-V4-Flash的智能体分别获得MRR 0.746和0.718,其差异不显著,但目标证据覆盖差距达0.561。结合发现22.1%至43.7%的命中前十成功案例仍存在至少一个图定义硬混淆因子排名更高,表明全池检索、硬混淆区分与可观察证据访问捕捉了模型行为的不同维度。GraphRareBench为更透明、证据感知的表型驱动诊断系统评估提供了基础。代码与数据见https://github.com/GUI0609/GraphRareBench。
原文摘要 · Abstract (English)
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternatives are ranked above it or what evidence a tool-using model examines before making its decision. We introduce GraphRareBench, a provenance-preserving benchmark containing 2,365 ontology-derived cases and 18,093 target-confounder pairs. Each case includes a coarsened HPO query, a fixed candidate pool, graph-defined hard confounders, and source-linked evidence records. On the 237-case gene-component-disjoint test split, supervised rankers using a shared 21-feature interface achieved MRRs ranging from 0.640 to 0.740 and case-averaged target-over-confounder accuracies ranging from 0.898 to 0.916. Agents instantiated with Agents-A1 and DeepSeek-V4-Flash achieved MRRs of 0.746 and 0.718, respectively. Their paired MRR difference was not statistically significant, whereas their target-evidence coverage differed by 0.561. Together with the observation that 22.1% to 43.7% of selected Hit@10 successes still ranked at least one graph-defined hard confounder above the target, these results indicate that full-pool retrieval, hard-confounder discrimination, and observable evidence access capture complementary aspects of model behavior. GraphRareBench therefore provides a foundation for more transparent and evidence-aware evaluation of phenotype-driven diagnostic systems. Code and data are available at https://github.com/GUI0609/GraphRareBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。