arXiv:2601.19773cs.CL2026-01ACL被引 2

诊断推理强不等于能有效问诊,新框架发现此差距并提出改进策略。

Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis

  • 用模拟患者和报告器建模问诊过程,量化信息覆盖程度。
  • 10个模型测试显示强推理无法保证有效收集证据,成性能瓶颈。
  • 提出REFINE策略,通过验证诊断主动消除不确定性,提升协作效率。

交互式医疗问诊要求智能体在不确定下主动获取缺失的临床证据。然而现有评估多为静态或结果导向,忽视证据收集过程。本文提出一种交互式评估框架,通过模拟患者和基于原子证据的模拟报告器显式建模咨询流程,并引入信息覆盖率(ICR)量化智能体在交互中揭示必要证据的完整性。为支持系统研究,构建EviMed基准,涵盖从常见症状到罕见疾病的多样化疾病场景,评估10种不同推理能力的模型。结果表明,强大的诊断推理并不等同于有效的信息采集,该不足是交互场景下性能受限的主要瓶颈。为此,提出REFINE策略,利用诊断验证引导智能体主动解决不确定性。大量实验表明,REFINE在多种数据集上持续优于基线,促进模型协作,使小型智能体在强推理监督下实现更优表现。代码已开源:https://github.com/NanshineLoong/EID-Benchmark。

原文摘要 · Abstract (English)

Interactive medical consultation requires an agent to proactively elicit missing clinical evidence under uncertainty. Yet existing evaluations largely remain static or outcome-centric, neglecting the evidence-gathering process. In this work, we propose an interactive evaluation framework that explicitly models the consultation process using a simulated patient and a \rev{simulated reporter} grounded in atomic evidences. Based on this representation, we introduce Information Coverage Rate (ICR) to quantify how completely an agent uncovers necessary evidence during interaction. To support systematic study, we build EviMed, an evidence-based benchmark spanning diverse conditions from common complaints to rare diseases, and evaluate 10 models with varying reasoning abilities. We find that strong diagnostic reasoning does not guarantee effective information collection, and this insufficiency acts as a primary bottleneck limiting performance in interactive settings. To address this, we propose REFINE, a strategy that leverages diagnostic verification to guide the agent in proactively resolving uncertainties. Extensive experiments demonstrate that REFINE consistently outperforms baselines across diverse datasets and facilitates effective model collaboration, enabling smaller agents to achieve superior performance under strong reasoning supervision. Our code can be found at https://github.com/NanshineLoong/EID-Benchmark .

医疗问答交互评估诊断推理信息采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。