用多模型集成提升医疗编码证据完整度,找回单模型遗漏的关键信息。
Complete Evidence Extraction with Model Ensembles: A Case Study on Medical Coding
- 通过融合多个语言模型的标记级证据,提升完整证据提取效果。
- 仅3个模型集成就超过最优单模型,召回率显著提升。
- 适合医疗合规、医保结算等需完整决策依据的高风险场景。
高风险决策依赖明确证据支持。现有研究多关注短而充分的证据,但监管合规与医疗计费要求完整的证据:所有支撑决策的相关输入标记。本文将完整证据提取定义为一项任务,并在医疗编码场景中开展案例研究。受Rashomon效应启发,我们聚合多个语言模型的标记级证据以提高证据完整性。基于性能相当的现有模型、特征归因方法及人工标注证据的数据集进行实验。结果表明,Rashomon集成可显著提升证据召回率,仅带来少量额外标记开销。仅三个模型的集成即超越最优单模型,恢复了个体模型遗漏的信息。
原文摘要 · Abstract (English)
High-stakes decisions informed by decision support systems require explicit evidence. While prior work focuses on short sufficient evidence, regulatory compliance and medical billing call for complete evidence: all relevant input tokens that support a decision. We formulate complete evidence extraction as a task and study it in a medical coding setting. Motivated by the Rashomon effect, we aggregate token-level evidence from multiple language models to increase evidence completeness. We perform a case study using existing equally-performing models, feature attributions, and a dataset with human-annotated evidence. Our results show that Rashomon ensembles significantly increase evidence recall while incurring only a small token overhead over individual models. Ensembles of only three models already outperform the best single model and recover information that individual models miss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。