arXiv:2503.02298cs.IR2025-03中稿 · Big Data Mining an…被引 2

用大模型零样本生成可解释的医生推荐,提升可信度与公平性。

A Zero-shot Explainable Doctor Ranking Framework with Large Language Models

  • 基于大模型动态生成疾病相关评分标准,实现零样本医生排序。
  • 在38种疾病-治疗组合上提升6.45 NDCG@10,优于最强基线。
  • 输出逐步推理过程,适合医疗推荐、可解释AI领域研究者使用。

在线医疗服务为患者提供了便捷的医生访问途径,但根据具体医疗需求有效排名医生仍具挑战。现有方法普遍缺乏对患者信任和知情决策至关重要的可解释性。此外,缺乏标准化基准和标注数据,限制了面向专业能力的医生排名发展。为此,我们提出一种基于大语言模型的零样本可解释医生排名框架。该框架动态生成疾病相关的评分标准,引导大模型以透明一致的方式评估医生相关性,并通过生成分步推理过程进一步增强可解释性。为支持严谨评估,我们构建并发布了DrRank——一个由38个疾病-治疗组合和4,325名医生资料组成的专家驱动数据集。在该基准上,我们的框架显著优于最强基线,NDCG@10提升6.45。全面分析表明,该框架在不同疾病类型、患者性别和地区间均表现公平。医学专家验证确认其结果可靠且可解释,具备实际应用潜力。此外,在BEIR基准的两个数据集上也验证了其广泛适用性。代码与数据已开源。

原文摘要 · Abstract (English)

Online medical service provides patients convenient access to doctors, but effectively ranking doctors based on specific medical needs remains challenging. Current ranking approaches typically lack the interpretability crucial for patient trust and informed decision-making. Additionally, the scarcity of standardized benchmarks and labeled data for supervised learning impedes progress in expertise-aware doctor ranking. To address these challenges, we propose an explainable ranking framework for doctor ranking powered by large language models in a zero-shot setting. Our framework dynamically generates disease-specific ranking criteria to guide the large language model in assessing doctor relevance with transparency and consistency. It further enhances interpretability by generating step-by-step rationales for its ranking decisions, improving the overall explainability of the information retrieval process. To support rigorous evaluation, we built and released DrRank, a novel expertise-driven dataset comprising 38 disease-treatment pairs and 4,325 doctor profiles. On this benchmark, our framework significantly outperforms the strongest baseline by +6.45 NDCG@10. Comprehensive analyses also show our framework is fair across disease types, patient gender, and geographic regions. Furthermore, verification by medical experts confirms the reliability and interpretability of our approach, reinforcing its potential for trustworthy, real-world doctor recommendation. To demonstrate its broader applicability, we validate our framework on two datasets from BEIR benchmark, where it again achieves superior performance. The code and associated data are available at: https://github.com/YangLab-BUPT/DrRank.

可解释性医生推荐大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。