arXiv:2608.18108cs.CLcs.AI2026-08中稿 · ICML

同一病情,不同上下文,模型分配医疗资源的决策竟截然相反。

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

论文配图:Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
图 1 · 摘自论文原文
  • 通过对比有无历史回答的上下文,测试模型决策变化
  • 三组模型中,新信息导致分配概率方向相反
  • 提醒在医疗决策中需警惕上下文对模型行为的影响

大型语言模型正被广泛应用于各类关键决策场景。尽管已有研究关注输入和情境框架带来的偏见,但模型在部署过程中累积的上下文也可能引发意外且不良的行为。本文以医疗资源分配为例,让模型在简短临床背景后,判断两人获得资源的概率;随后在同一情景中加入一句包含相反患者信息的额外句子,考察该信息是否包含此前模型的回答。在四组测试模型中的三组,当新信息出现时,模型概率分布发生显著变化,且方向常相反(如支持人物B vs. 支持人物A)。我们进一步通过多维度属性调整验证了这一现象。结果表明,在敏感医疗场景中,患者信息的上下文依赖性会显著影响模型行为。本工作强调在决策系统中谨慎集成大模型、进行上下文工程及行为研究的重要性。

原文摘要 · Abstract (English)

Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical context, and then sees the same scenario with a single extra sentence containing contrasting patient information, either with or without its previous response in context. Across three of four tested models, the paired-context and independent-inference experiments have different probability shifts, often in opposite directions (in favor of Person B vs. in favor of Person A) when new information is provided. We include additional paired-context experiments to show the effect of varying attributes across scenario axes. Our findings show the context-dependent effect of patient information in a sensitive medical use case. More broadly, our work shows the importance of carefully incorporating LLM-based systems into decision-making processes, context engineering, and further model behavioral studies.

医疗决策模型行为上下文依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。