arXiv:2603.21460cs.IRcs.AI2026-03

用AI比对移植手册差异,发现内容不统一问题严重

When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models

  • 用检索增强模型把同一问题映射到不同医院手册
  • 20.8%的对比显示临床意义差异,生殖健康缺失超95%
  • 适合医疗政策研究者和指南制定者参考

美国各移植中心的患者教育材料差异显著,但缺乏系统量化方法。本文提出基于检索增强语言模型的框架,将相同患者问题映射至不同中心的手册,并通过五标签一致性分类体系比较答案。在23家中心的102份手册与1,115个基准问题上应用,从问题、主题、器官和中心四个维度量化异质性。结果发现20.8%的非空对比存在临床意义分歧,集中于病情监测和生活方式主题;覆盖缺口更严重:96.2%的问题-手册对缺失相关内容,生殖健康缺失率达95.1%。中心级差异模式稳定可解释,反映系统性机构差异,可能源于患者多样性。该研究揭示了移植患者教育材料的信息鸿沟,为内容优化提供了依据。

原文摘要 · Abstract (English)

Patient education materials for solid-organ transplantation vary substantially across U.S. centers, yet no systematic method exists to quantify this heterogeneity at scale. We introduce a framework that grounds the same patient questions in different centers' handbooks using retrieval-augmented language models and compares the resulting answers using a five-label consistency taxonomy. Applied to 102 handbooks from 23 centers and 1,115 benchmark questions, the framework quantifies heterogeneity across four dimensions: question, topic, organ, and center. We find that 20.8% of non-absent pairwise comparisons exhibit clinically meaningful divergence, concentrated in condition monitoring and lifestyle topics. Coverage gaps are even more prominent: 96.2% of question-handbook pairs miss relevant content, with reproductive health at 95.1% absence. Center-level divergence profiles are stable and interpretable, where heterogeneity reflects systematic institutional differences, likely due to patient diversity. These findings expose an information gap in transplant patient education materials, with document-grounded medical question answering highlighting opportunities for content improvement.

医疗AI文档对比知识差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。