arXiv:2511.17818cs.LGcs.AI2025-11

用大模型生成医疗反事实数据,提升政策评估准确性。

APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs

  • 用大模型预测不同治疗下临床特征变化,生成反事实标注
  • 在MIMIC-IV上验证,大幅改善策略评估结果
  • 提供熵指标判断标注增益何时停止,适合医疗决策研究者

离策略评估(OPE)可在部署前估算上下文无关策略的价值,对医疗等高风险领域至关重要。但传统OPE受限于行为数据集的规模与覆盖范围。以往方法依赖专家标注的反事实数据,成本高昂。本文提出利用大语言模型(LLMs)生成医疗领域的反事实标注:基于领域知识引导LLM预测关键临床特征在不同治疗下的演变,并通过已知奖励函数转换为反事实数据。我们在MIMIC-IV的两个患者子集中评估多个LLM预测临床特征的能力,发现当前最优模型表现相当。在此基础上,我们生成基于LLM的反事实标注并集成至OPE估计器。实证结果显示,在行为策略与目标策略存在不同程度偏移时,这些标注显著提升评估精度,直至达到信息饱和点。我们提出一种基于熵的度量来识别该临界点。结果表明,基于大模型的反事实标注是应对医疗数据覆盖不足的可扩展方案,有助于更安全地部署临床决策策略。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) estimates the value of a contextual bandit policy prior to deployment. As such, OPE plays a critical role in ensuring safety in high-stakes domains such as healthcare. However, standard OPE approaches are limited by the size and coverage of the behavior dataset. While previous work has explored using expert-labeled counterfactual annotations to enhance dataset coverage, obtaining such annotations is expensive, limiting the scalability of prior approaches. We propose leveraging large language models (LLMs) to generate counterfactual annotations for OPE in medical domains. Our method uses domain knowledge to guide LLMs in predicting how key clinical features evolve under alternate treatments. These predicted features can then be transformed using known reward functions to create counterfactual annotations. We first evaluate the ability of several LLMs to predict clinical features across two patient subsets in MIMIC-IV, finding that state-of-the-art LLMs achieve comparable performance. Building on this capacity to predict clinical features, we generate LLM-based counterfactual annotations and incorporate them into an OPE estimator. Our empirical results analyze the benefits of counterfactual annotations under varying degrees of shift between the behavior and target policies. We find that in most cases, the LLM-based counterfactual annotations significantly improve OPE estimates up to a point. We provide an entropy-based metric to identify when additional annotations cease to be useful. Our results demonstrate that LLM-based counterfactual annotations offer a scalable approach for addressing coverage limitations in healthcare datasets, enabling safer deployment of decision-making policies in clinical settings.

策略评估大模型应用医疗决策反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。