针对低收入国家医疗场景的专用检索增强系统,性能超越前沿大模型。
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
- 基于印度本土医疗数据构建专属知识库,实现精准临床决策支持。
- 在4023道题中得分51.9%,领先GPT-5.4等主流模型。
- 适合关注医疗AI落地、本地化系统设计的研究者与开发者。
通用大语言模型近期在医学基准上表现优异,但评估样本有限且多基于高收入国家数据。本文评估了专为印度及中低收入国家(LMIC)医疗场景设计的检索增强生成(RAG)系统VITA。VITA从疾病指南、印度抗菌药物耐药性数据、国家药品目录约束和资源受限护理协议等定制语料库中检索信息;其架构与语料库为专有,但基准、医生撰写的评分标准及完整响应与评分结果公开可验证。在4,023道英文HealthBench问题(占基准80.5%)上,由GPT-4.1评分,VITA以51.9%的满分得分排名第一,优于GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),并在45.4%的问题上取得最高分。为测试对新模型的鲁棒性,500题子集再次测试当前模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3),由无关联的开源模型DeepSeek-V4-Pro评分。此时差距缩小至无统计差异:VITA与GPT-5.5平均得分相近,但VITA在加权得分和胜出问题数上领先。其准确性与完整性优势在中立评分下依然存在,但沟通表达得分较低。结果表明,专用临床RAG系统在开放基准上仍可媲美前沿大模型,体现语料库特异性作为设计变量能提升知识锚定能力,代价是沟通流畅度下降。
原文摘要 · Abstract (English)
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。