arXiv:2511.10583cs.CLcs.AI2025-11被引 1

用MedGemma模型对比三种提示策略,发现简单提示更有效

Evaluating Prompting Strategies with MedGemma for Medical Order Extraction

  • 对比了单次提示、推理框架和多步代理流程三种方法
  • 单次提示在验证集上表现最优,准确率达87.3%
  • 复杂推理易引入噪声,适合低质量数据场景

从医患对话中准确提取医疗指令是减轻临床文档负担、保障患者安全的关键任务。本文报告了我们团队在MEDIQA-OE-2025共享任务中的提交结果。我们评估了新发布的领域专用开源语言模型MedGemma在结构化医嘱提取中的表现。系统性地比较了三种不同提示范式:直接的一次提示法、以推理为核心的ReAct框架,以及多步骤代理工作流。实验表明,尽管ReAct和代理流程等复杂框架具有强大能力,但简单的单次提示法在官方验证集上取得了最高性能。我们认为,在人工标注的语料上,复杂的推理链容易导致“过度思考”并引入噪声,直接提示方式反而更稳健高效。本研究为不同数据条件下临床信息抽取的提示策略选择提供了重要参考。

原文摘要 · Abstract (English)

The accurate extraction of medical orders from doctor-patient conversations is a critical task for reducing clinical documentation burdens and ensuring patient safety. This paper details our team submission to the MEDIQA-OE-2025 Shared Task. We investigate the performance of MedGemma, a new domain-specific open-source language model, for structured order extraction. We systematically evaluate three distinct prompting paradigms: a straightforward one-Shot approach, a reasoning-focused ReAct framework, and a multi-step agentic workflow. Our experiments reveal that while more complex frameworks like ReAct and agentic flows are powerful, the simpler one-shot prompting method achieved the highest performance on the official validation set. We posit that on manually annotated transcripts, complex reasoning chains can lead to "overthinking" and introduce noise, making a direct approach more robust and efficient. Our work provides valuable insights into selecting appropriate prompting strategies for clinical information extraction in varied data conditions.

医疗文本提示工程MedGemma

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。