首个公开的医疗编码模型评测基准,专测大模型对医保分组逻辑的理解与推理能力。
The NordDRG AI Benchmark for Large Language Models
- 构建可机器读取的完整医保分组规则库和治理流程文档
- 在13个逻辑任务中顶级模型全对,但全链路模拟时仅部分模型通过
- 适合医疗AI评估、医保系统开发及政策透明度研究者使用
大型语言模型正被用于临床编码与决策支持,但尚无公开基准针对决定医院报销金额的诊断相关分组(DRGs)层面。在多数OECD国家,DRGs通过受控的分组软件分配占数万亿美元规模的医疗支出,透明性与可审计性至关重要。我们发布NordDRG-AI-Benchmark,首个公开、规则完整的DRG推理测试平台。包含(i)约20页可机器读取的NordDRG定义表,(ii)专家手册与变更日志模板,涵盖治理工作流。平台提供两个套件:13项逻辑任务(代码查找、跨表推理、分组特征识别、多语言术语处理、CC/MCC有效性校验)与13项分组器任务,要求完全模拟分组器并以精确匹配评分DRG及触发dr_glogic.id。轻量级参考代理(LogicAgent、GrouperAgent)支持仅基于产物的评估。在无网络的产物仅评估环境下,GPT-5 Thinking和Opus 4.1在13项逻辑任务中均获13/13,o3得12/13;中端模型(GPT-5 Thinking Mini, o4-mini, GPT-5 Fast)得6-8/13,其余模型均≤5/13。在13项全链路分组器模拟中,GPT-5 Thinking解决7/13,o3 6/13,o4-mini 3/13;GPT-5 Thinking Mini仅1/13,其余全部0/13。据我们所知,这是首次公开报告大模型部分模拟完整NordDRG分组逻辑并具备治理级可追溯性的成果。结合规则完备发布、精确匹配任务与开放评分,为医院资金领域的横向与纵向评估提供可复现标尺。基准材料已开源于Github。
原文摘要 · Abstract (English)
Large language models (LLMs) are being piloted for clinical coding and decision support, yet no open benchmark targets the hospital-funding layer where Diagnosis-Related Groups (DRGs) determine reimbursement. In most OECD systems, DRGs route a substantial share of multi-trillion-dollar health spending through governed grouper software, making transparency and auditability first-order concerns. We release NordDRG-AI-Benchmark, the first public, rule-complete test bed for DRG reasoning. The package includes (i) machine-readable approximately 20-sheet NordDRG definition tables and (ii) expert manuals and change-log templates that capture governance workflows. It exposes two suites: a 13-task Logic benchmark (code lookup, cross-table inference, grouping features, multilingual terminology, and CC/MCC validity checks) and a 13-task Grouper benchmark that requires full DRG grouper emulation with strict exact-match scoring on both the DRG and the triggering drg_logic.id. Lightweight reference agents (LogicAgent, GrouperAgent) enable artefact-only evaluation. Under an artefact-only (no web) setting, on the 13 Logic tasks GPT-5 Thinking and Opus 4.1 score 13/13, o3 scores 12/13; mid-tier models (GPT-5 Thinking Mini, o4-mini, GPT-5 Fast) achieve 6-8/13, and remaining models score 5/13 or below. On full grouper emulation across 13 tasks, GPT-5 Thinking solves 7/13, o3 6/13, o4-mini 3/13; GPT-5 Thinking Mini solves 1/13, and all other tested endpoints score 0/13. To our knowledge, this is the first public report of an LLM partially emulating the complete NordDRG grouper logic with governance-grade traceability. Coupling a rule-complete release with exact-match tasks and open scoring provides a reproducible yardstick for head-to-head and longitudinal evaluation in hospital funding. Benchmark materials available in Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。