arXiv:2512.03737cs.CLcs.IR2025-12

用大模型提升医疗搜索相关性,避免幻觉且效果显著

AR-Med: Automated Relevance Enhancement in Medical Search via LLM-Driven Information Augmentation

  • 通过检索增强让大模型基于真实医学知识推理
  • 离线准确率达93%,比原系统高24个百分点
  • 适合需要可靠医疗搜索的平台与开发者

在线医疗平台的精准搜索对用户安全和服务效能至关重要。传统方法难以理解复杂查询,而大语言模型虽具强大语义理解能力,但在医疗领域面临事实幻觉、专业知识不足和高成本等挑战。为此,我们提出AR-Med框架,通过检索增强机制将大模型推理锚定在可信医学知识上,确保高准确率与可靠性。为支持高效线上服务,设计了知识蒸馏方案,将大模型压缩为轻量级学生模型。同时构建了LocalQSMed多专家标注基准,指导模型迭代并保障离线与线上性能一致。大量实验表明,AR-Med离线准确率超过93%,较原始在线系统提升24个百分点,并显著提高线上相关性与用户满意度。本工作为真实医疗场景中可信大模型系统的构建提供了可落地、可扩展的范式。

原文摘要 · Abstract (English)

Accurate and reliable search on online healthcare platforms is critical for user safety and service efficacy. Traditional methods, however, often fail to comprehend complex and nuanced user queries, limiting their effectiveness. Large language models (LLMs) present a promising solution, offering powerful semantic understanding to bridge this gap. Despite their potential, deploying LLMs in this high-stakes domain is fraught with challenges, including factual hallucinations, specialized knowledge gaps, and high operational costs. To overcome these barriers, we introduce \textbf{AR-Med}, a novel framework for \textbf{A}utomated \textbf{R}elevance assessment for \textbf{Med}ical search that has been successfully deployed at scale on the Online Medical Delivery Platforms. AR-Med grounds LLM reasoning in verified medical knowledge through a retrieval-augmented approach, ensuring high accuracy and reliability. To enable efficient online service, we design a practical knowledge distillation scheme that compresses large teacher models into compact yet powerful student models. We also introduce LocalQSMed, a multi-expert annotated benchmark developed to guide model iteration and ensure strong alignment between offline and online performance. Extensive experiments show AR-Med achieves an offline accuracy of over 93\%, a 24\% absolute improvement over the original online system, and delivers significant gains in online relevance and user satisfaction. Our work presents a practical and scalable blueprint for developing trustworthy, LLM-powered systems in real-world healthcare applications.

医疗搜索大模型知识增强模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。