arXiv:2512.18999cs.CLcs.AI2025-12

用模块化设计提升大模型在医疗随访中的对话稳定性和信息提取准确率

Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework

  • 将随访任务分解为多个子任务,通过结构化流程控制对话流
  • 对话轮次减少46.73%,令牌消耗降低80%至87.5%,信息提取更准确
  • 适合医疗场景中对可靠性要求高的大模型应用部署

在直接以端到端方式应用于医疗随访任务时,大型语言模型(LLMs)常因随访表单的复杂性而出现对话流程失控和信息提取不准确的问题。为此,我们设计并比较了两种随访聊天机器人系统:基于端到端大模型的对照组与采用结构化流程控制的模块化流水线实验组。实验结果表明,尽管端到端方法在长且复杂的表单上频繁失败,但我们的模块化方法——基于任务分解、语义聚类和流程管理——显著提升了对话稳定性和信息提取准确性。此外,该方法将对话轮次减少46.73%,令牌消耗降低80%至87.5%。这些发现强调了在高风险医疗随访场景中部署大模型时,整合外部控制机制的必要性。

原文摘要 · Abstract (English)

When applied directly in an end-to-end manner to medical follow-up tasks, Large Language Models (LLMs) often suffer from uncontrolled dialog flow and inaccurate information extraction due to the complexity of follow-up forms. To address this limitation, we designed and compared two follow-up chatbot systems: an end-to-end LLM-based system (control group) and a modular pipeline with structured process control (experimental group). Experimental results show that while the end-to-end approach frequently fails on lengthy and complex forms, our modular method-built on task decomposition, semantic clustering, and flow management-substantially improves dialog stability and extraction accuracy. Moreover, it reduces the number of dialogue turns by 46.73% and lowers token consumption by 80% to 87.5%. These findings highlight the necessity of integrating external control mechanisms when deploying LLMs in high-stakes medical follow-up scenarios.

医疗随访大模型应用对话系统流程控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。