arXiv:2604.22067cs.CLcs.AI2026-04

用AI智能选问题,提升精神科问诊信息获取效率。

Optimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake

论文配图:Optimal Question Selection from a Large Question Bank for Clinical Field Recovery in Conversational Psychiatric Intake
图 1 · 摘自论文原文
  • 基于655个临床问题库,设计自适应提问策略。
  • 大模型引导策略在5种行为条件下信息恢复率最高。
  • 尤其对沉默寡言患者,动态提问比固定流程更有效。

精神科问诊是高风险、序列化的信息采集过程,临床医生需在时间有限下决定问什么、何时问及如何解读不完整或模糊的回答。尽管医疗对话AI日益受关注,但该场景仍缺乏基础设施支持。本文将此任务建模为基于临床问题库的问答选择问题,设定目标信息并控制患者行为难度。我们构建了一个包含655个临床医师编写的问题和对应5种行为状态的合成患者案例的基准测试。评估涵盖300次访谈,涉及4名患者与5种行为条件,对比随机提问、固定问诊表与大模型引导的自适应策略。结果表明,固定问诊表显著优于随机提问,而大模型引导策略在整体表现上最优;在防御性或简略型患者中,自适应策略的优势尤为明显。研究揭示,对话式临床系统的表现不仅取决于信息披露后的理解能力,更关键在于能否在有限交互预算内触及关键话题。该基准为研究临床结构与动态追问如何促进信息恢复提供了可控框架。

原文摘要 · Abstract (English)

Psychiatric intake is a sequential, high-stakes information-gathering process in which clinicians must decide what to ask, in what order, and how to interpret incomplete or ambiguous responses under limited time. Despite growing interest in conversational AI for healthcare, there is still limited infrastructure for conversational AI in this application. Accordingly, we formulate this task as a question-selection problem with clinically grounded questions, known target information, and controllable patient difficulty. We also introduce a task-specific question-selection benchmark based on a bank of 655 clinician-authored intake questions and corresponding synthetic patient vignettes with 5 different behavioral conditions. In our evaluation, we compare random questioning, a clinical psychiatric intake form baseline, and an LLM-guided adaptive policy across 300 interview sessions spanning four patients and five behavioral conditions. Across the benchmark, the clinically ordered fixed form substantially outperforms random questioning, and the LLM-guided policy achieves the strongest overall recovery. The advantage of adaptation grows sharply under patient behavior that is less amenable to field recovery, especially under guarded-concise conditions. These findings suggest that performance in conversational clinical systems depends not only on language understanding after information is disclosed, but also on whether the system reaches the right topics within a limited interaction budget. More broadly, the benchmark provides a controlled framework for studying how clinical structure and adaptive follow-up contribute to information recovery in interactive clinical machine learning.

精神科问诊对话AI自适应提问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。