测试大模型在药品剂量时间不确定下的决策能力
Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA

- 构建81个真实用药场景,评估模型对24小时滚动剂量的推理能力
- 模型在时间窗口计算和模糊情况处理上错误率高,且自信回应可能出错
- 适合关注医疗问答安全性和时序推理的开发者与研究者
大型语言模型(LLMs)被越来越多用于日常健康咨询,包括判断用户是否可安全服用另一剂非处方药(OTC)。然而,此类关键安全场景在现有医学问答评估中仍被忽视。正确回答需追踪服药时间、计算24小时内累计摄入量、遵守产品标签限制,并处理不完整用药史。我们提出DOSEBENCH,一个包含81个精心设计的成人对乙酰氨基酚和布洛芬用药场景的基准,配有手工标注的参考答案。我们评估了四种LLMs在多次运行中的表现,使用决策正确性、一致性、解释可验证性、错误类型及置信度信号等指标,共生成1,620条模型响应。结果表明,模型在滚动窗口推理和敏感模糊情境中频繁出错,且看似稳定或自信的回复仍可能违反剂量约束。这些发现表明,非处方药剂量问答为评估医疗问答中的时序推理、规则遵循与安全不确定性处理提供了精准而实用的测试平台。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the-counter (OTC) medication. Yet this common safety-relevant setting remains underexplored in existing medical QA evaluations, where correct answers require tracking dose timing, computing rolling 24-hour intake, following product-label constraints, and handling incomplete medication histories. We introduce DOSEBENCH, a focused benchmark of 81 curated OTC dosing scenarios focused on adult acetaminophen and ibuprofen use, with manually annotated gold references. We evaluate four LLMs across repeated runs using metrics for decision correctness, consistency, explanation verifiability, failure types, and confidence-related signals, resulting in 1,620 model responses. Our results show that models frequently struggle with rolling-window reasoning and ambiguity-sensitive cases and that stable or confident-looking responses can still violate dosing constraints. These findings suggest that OTC dosing QA provides a narrow yet practical testbed for evaluating temporal reasoning, constraint following, and safety-relevant uncertainty handling in medical QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。