把医生拒绝对话模型的决策当作隐式偏好信号,提升AI在价值医疗中的实用性。
Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care
- 将医生拒绝视为专家级偏好信号,结合患者状态与医生能力建模更新目标。
- 提出双学习架构,避免因医生能力不足导致正确建议被系统压制。
- 适用于需长期追踪疗效的价值医疗场景,适合临床决策支持系统优化者。
我们将临床医生对AI建议的拒绝对话重新定义为隐式偏好数据——与强化学习中的人类反馈机制结构相同,但更丰富:标注者是领域专家,备选方案具有真实后果,下游结果可观测。本文提出一个形式化框架,扩展标准偏好学习:构建五类拒绝对话分类体系,映射至不同模型更新目标;提出依赖患者状态s、组织背景c和医生能力κ的偏好公式,其中κ分解为执行能力κ-exec与对齐能力κ-align;设计双学习架构,通过交替优化联合训练奖励模型与能力模型,防止一种称为抑制偏误的失效模式——当医生执行能力低于阈值时,系统会系统性压制正确但困难的建议。我们认为,在基于结果支付的慢性病管理中,拒绝对话数据具备独特优势:纵向数据密集、决策空间集中、结果标签可得、能力差异自然存在;而结合纵向结果测量与对齐激励的训练环境,是学习与患者轨迹一致而非仅匹配就诊经济的奖励模型的必要条件。该框架源自一次真实价值医疗部署中提升医生能力的实操工作。
原文摘要 · Abstract (English)
We reframe clinician overrides of clinical AI recommendations as implicit preference data - the same signal structure exploited by reinforcement learning from human feedback (RLHF), but richer: the annotator is a domain expert, the alternatives carry real consequences, and downstream outcomes are observable. We present a formal framework extending standard preference learning with three contributions: a five-category override taxonomy mapping override types to distinct model update targets; a preference formulation conditioned on patient state s, organizational context c, and clinician capability kappa, where kappa decomposes into execution capability kappa-exec and alignment capability kappa-align; and a dual learning architecture that jointly trains a reward model and a capability model via alternating optimization, preventing a failure mode we term suppression bias-the systematic suppression of correct-but-difficult recommendations when clinician capability falls below the execution threshold. We argue that chronic disease management under outcome-based payment contracts produces override data with uniquely favorable properties-longitudinal density, concentrated decision space, outcome labels, and natural capability variation-and that training environments combining longitudinal outcome measurement with aligned financial incentives are a necessary condition for learning a reward model aligned with patient trajectory rather than with encounter economics. This framework emerged from operational work to improve clinician capability in a live value-based care deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。