arXiv:2604.08788cs.CL2026-04被引 1

构建医疗对话中隐藏关切推理的评估基准,模拟真实诊疗中的信息不对称。

MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

论文配图:MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability
图 1 · 摘自论文原文
  • 设计交互式患者模拟器,隐藏真实关切以测试模型挖掘能力。
  • 300个病例、600次人机交互,验证模型在多轮对话中确认隐性关切的能力。
  • 适合研究医疗对话系统、临床决策支持与人机交互的学者使用。

患者与医生之间的沟通存在信息不对称:患者常不主动披露担忧、误解或实际障碍,除非医生巧妙引导。有效医疗对话需在部分可观测条件下进行推理——医生须通过互动揭示潜在关切,并以合适方式引导患者接受适当治疗。然而,现有医疗对话基准大多忽略这一挑战,或直接暴露隐藏状态,将引导过程简化为信息提取,或未建模隐藏内容即评价回复。我们提出MedConceal,一个包含交互式患者模拟器的基准,用于评估医疗对话中的隐藏关切推理,涵盖300个精心筛选的案例和600次医生-大模型交互。案例源自临床医生解答的在线健康讨论,每个案例包含医生可见背景与模拟器内部隐藏关切,后者基于文献并由专家制定分类体系生成。模拟器向对话代理隐藏这些关切,通过理论基础的回合级通信信号追踪其是否被揭示与解决,并经医生评审确保临床合理性。该设计实现对任务成功及互动过程的双重评估。我们考察两种能力:确认(通过多轮对话揭示隐藏关切)与干预(解决主要问题并引导患者走向目标方案)。结果表明,无单一系统全面领先:前沿模型在不同确认指标上表现突出,而人类医生(N=159)在干预成功率上仍最优。综合结果表明,在部分可观测条件下推理隐藏关切是医疗对话系统的关键未解难题。

原文摘要 · Abstract (English)

Patient-clinician communication is an asymmetric-information problem: patients often do not disclose fears, misconceptions, or practical barriers unless clinicians elicit them skillfully. Effective medical dialogue therefore requires reasoning under partial observability: clinicians must elicit latent concerns, confirm them through interaction, and respond in ways that guide patients toward appropriate care. However, existing medical dialogue benchmarks largely sidestep this challenge by exposing hidden patient state, collapsing elicitation into extraction, or evaluating responses without modeling what remains hidden. We present MedConceal, a benchmark with an interactive patient simulator for evaluating hidden-concern reasoning in medical dialogue, comprising 300 curated cases and 600 clinician-LLM interactions. Built from clinician-answered online health discussions, each case pairing clinician-visible context with simulator-internal hidden concerns derived from prior literature and structured using an expert-developed taxonomy. The simulator withholds these concerns from the dialogue agent, tracks whether they have been revealed and addressed via theory-grounded turn-level communication signals, and is clinician-reviewed for clinical plausibility. This enables process-aware evaluation of both task success and the interaction process that leads to it. We study two abilities: confirmation, surfacing hidden concerns through multi-turn dialogue, and intervention, addressing the primary concern and guiding the patient toward a target plan. Results show that no single system dominates: frontier models lead on different confirmation metrics, while human clinicians (N=159) remain strongest on intervention success. Together, these results identify hidden-concern reasoning under partial observability as a key unresolved challenge for medical dialogue systems.

医疗对话隐藏关切评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。