arXiv:2506.21532cs.CLcs.AI2025-06EMNLP被引 8

分析1.1万条真实医患对话,揭示LLM问诊中的风险与改进方向

"What's Up, Doc?": Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasets

  • 从海量对话中筛选出1.1万条真实医疗咨询数据,构建健康问答专用数据集
  • 发现用户常因信息不全提问,且存在诱导性问题导致模型盲目迎合
  • 为医疗类大模型的交互设计与安全优化提供实证依据,适合医疗AI研究者

人们越来越多地通过交互式聊天机器人向大语言模型(LLMs)寻求医疗信息,但这些对话的性质及潜在风险仍不明确。本文通过筛选大规模对话数据集,构建了包含11,000条真实对话、25,000条用户消息的HealthChat-11K数据集。利用该数据集与临床医生制定的分类体系,我们系统分析了用户在21个不同医学专科中寻求健康信息的行为模式。结果揭示了常见交互形式、上下文缺失现象、情绪化表达以及可能诱发模型讨好行为的引导性提问,凸显了当前部署于对话AI中的大语言模型在医疗支持能力上的不足。相关代码与数据可访问:https://github.com/yahskapar/HealthChat

原文摘要 · Abstract (English)

People are increasingly seeking healthcare information from large language models (LLMs) via interactive chatbots, yet the nature and inherent risks of these conversations remain largely unexplored. In this paper, we filter large-scale conversational AI datasets to achieve HealthChat-11K, a curated dataset of 11K real-world conversations composed of 25K user messages. We use HealthChat-11K and a clinician-driven taxonomy for how users interact with LLMs when seeking healthcare information in order to systematically study user interactions across 21 distinct health specialties. Our analysis reveals insights into the nature of how and why users seek health information, such as common interactions, instances of incomplete context, affective behaviors, and interactions (e.g., leading questions) that can induce sycophancy, underscoring the need for improvements in the healthcare support capabilities of LLMs deployed as conversational AI. Code and artifacts to retrieve our analyses and combine them into a curated dataset can be found here: https://github.com/yahskapar/HealthChat

医疗AI对话分析大模型风险数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。