小模型Med-V1实现低成本精准医学证据溯源,效果媲美大模型。
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution
- 用三亿参数小模型+自研合成数据训练,提升证据验证准确率27%~71%
- 在五项生物医学任务上表现接近GPT-5,且生成解释质量高
- 可检测临床指南中的错误引用,适合医疗合规与大模型纠错场景
判断一篇文章是否支持某个断言,对幻觉检测和声明验证至关重要。尽管大型语言模型(LLMs)具备自动化该任务的潜力,但要达到良好性能需依赖像GPT-5这样的前沿模型,其部署成本过高。为高效完成生物医学证据溯源,我们提出Med-V1,一个仅含三亿参数的小型语言模型家族。该模型在本研究新构建的高质量合成数据上训练,其在五个统一为验证格式的生物医学基准上表现显著优于基线模型(提升27.0%至71.3%)。尽管规模较小,Med-V1性能可比肩如GPT-5等前沿大模型,并能生成高质量预测解释。我们首次使用Med-V1开展案例研究,量化不同引文指令下大模型生成答案的幻觉情况:结果显示引文格式指令显著影响引文有效性与幻觉率,其中GPT-5虽生成更多主张,但幻觉率与GPT-4o相当。此外,第二项案例研究显示,Med-V1可自动识别临床实践指南中高风险的证据误引,揭示可能带来负面公共健康影响的问题,此类问题传统方法难以规模化发现。总体而言,Med-V1为生物医学证据溯源与验证任务提供了高效、准确的轻量级替代方案。Med-V1代码已公开于https://github.com/ncbi-nlp/Med-V1。
原文摘要 · Abstract (English)
Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language models (LLMs) have the potential to automate this task, achieving strong performance requires frontier models such as GPT-5 that are prohibitively expensive to deploy at scale. To efficiently perform biomedical evidence attribution, we present Med-V1, a family of small language models with only three billion parameters. Trained on high-quality synthetic data newly developed in this study, Med-V1 substantially outperforms (+27.0% to +71.3%) its base models on five biomedical benchmarks unified into a verification format. Despite its smaller size, Med-V1 performs comparably to frontier LLMs such as GPT-5, along with high-quality explanations for its predictions. We use Med-V1 to conduct a first-of-its-kind use case study that quantifies hallucinations in LLM-generated answers under different citation instructions. Results show that the format instruction strongly affects citation validity and hallucination, with GPT-5 generating more claims but exhibiting hallucination rates similar to GPT-4o. Additionally, we present a second use case showing that Med-V1 can automatically identify high-stakes evidence misattributions in clinical practice guidelines, revealing potentially negative public health impacts that are otherwise challenging to identify at scale. Overall, Med-V1 provides an efficient and accurate lightweight alternative to frontier LLMs for practical and real-world applications in biomedical evidence attribution and verification tasks. Med-V1 is available at https://github.com/ncbi-nlp/Med-V1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。