arXiv:2601.11544cs.HCcs.AI2026-01中稿 · 2025 IEEE Internat…被引 1

用大模型做用药咨询,既要灵活又要可靠。

Medication counseling with large language models: balancing flexibility and rigidity

  • 聚焦长期用药对话,通过约束机制提升稳定性
  • 在保持自然对话能力的同时减少幻觉和错误
  • 适合医疗对话系统开发者与临床辅助研究者

大型语言模型(LLM)显著提升了软件代理的交互能力,使其能以类人方式灵活沟通。然而,在药房等高风险场景中,过度灵活易引发错误;过于僵化则无法应对未预设情境。现有研究多关注宽泛、简短的药物咨询,而本文反其道而行之,聚焦狭窄但持续时间长的用药咨询任务。这不仅深化了对复杂对话的理解,也揭示了长期交互中的挑战。核心难题仍在于如何平衡对话规范性与灵活性。为此,我们提出一个原型系统,旨在实现安全可靠的用药咨询。设计重点包括满足对话要求、降低幻觉、生成高质量回应的方法。这些方法可在增强系统确定性的同时,保留LLM带来的动态交互能力。但该系统仍需持续测试与人工介入,且应在非标准基准下评估,因常规评测难以反映此类系统的复杂性。

原文摘要 · Abstract (English)

The introduction of large language models (LLMs) has greatly enhanced the capabilities of software agents. Instead of relying on rule-based interactions, agents can now interact in flexible ways akin to humans. However, this flexibility quickly becomes a problem in fields where errors can be disastrous, such as in a pharmacy context, but the opposite also holds true; a system that is too inflexible will also lead to errors, as it can become too rigid to handle situations that are not accounted for. Work using LLMs in a pharmacy context have adopted a wide scope, accounting for many different medications in brief interactions -- our strategy is the opposite: focus on a more narrow and long task. This not only enables a greater understanding of the task at hand, but also provides insight into what challenges are present in an interaction of longer nature. The main challenge, however, remains the same for a narrow and wide system: it needs to strike a balance between adherence to conversational requirements and flexibility. In an effort to strike such a balance, we present a prototype system meant to provide medication counseling while juggling these two extremes. We also cover our design in constructing such a system, with a focus on methods aiming to fulfill conversation requirements, reduce hallucinations and promote high-quality responses. The methods used have the potential to increase the determinism of the system, while simultaneously not removing the dynamic conversational abilities granted by the usage of LLMs. However, a great deal of work remains ahead, and the development of this kind of system needs to involve continuous testing and a human-in-the-loop. It should also be evaluated outside of commonly used benchmarks for LLMs, as these do not adequately capture the complexities of this kind of conversational system.

用药咨询大模型对话系统医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。