提出医学对话基准MedDialBench,量化患者行为对大模型诊断鲁棒性的影响。
MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors

- 分解患者行为为五维,分级设计并控制交互条件
- 发现虚构症状导致的诊断误差是隐瞒信息的1.7至3.4倍
- 仅虚构行为引发跨维度超加性效应,且不可通过追问弥补
交互式医疗对话基准显示,当面对不合作患者时,大语言模型(LLM)诊断准确率显著下降。现有方法或未分级评估、或缺乏案例具体性,或仅简化为单一维度,且均未分析多维度间交互。本文提出MedDialBench,支持对患者行为各维度在剂量-响应关系下的系统评估。该基准将患者行为拆解为五个维度——逻辑一致性、健康认知、表达风格、披露程度与态度,每个维度含分级严重度及案例特异性脚本。基于此控制因子设计,实现敏感性分析、剂量反应建模及跨维度交互检测。在7,225次对话(85个病例 × 17种配置 × 5个模型)中评估五款前沿模型,发现根本不对称:信息污染(虚构症状)造成的准确率下降是信息缺失(隐瞒信息)的1.7–3.4倍;且仅虚构配置在所有五模型中均达统计显著(McNemar p < 0.05)。六组维度组合中,仅涉及虚构的行为组合产生超加性效应:所有含虚构组合的观测/期望比(O/E)为0.70–0.81(35%–44%的可达成病例因组合失败,尽管单独维度下成功),而无虚构组合均为纯加性效应(O/E ≈ 1.0)。提问策略可缓解信息缺失但无法抵消信息污染;模型漏洞各异,最差情况下准确率下降达38.8至54.1个百分点。
原文摘要 · Abstract (English)
Interactive medical dialogue benchmarks have shown that LLM diagnostic accuracy degrades significantly when interacting with non-cooperative patients, yet existing approaches either apply adversarial behaviors without graded severity or case-specific grounding, or reduce patient non-cooperation to a single ungraded axis, and none analyze cross-dimension interactions. We introduce MedDialBench, a benchmark enabling controlled, dose-response characterization of how individual patient behavior dimensions affect LLM diagnostic robustness. It decomposes patient behavior into five dimensions -- Logic Consistency, Health Cognition, Expression Style, Disclosure, and Attitude -- each with graded severity levels and case-specific behavioral scripts. This controlled factorial design enables graded sensitivity analysis, dose-response profiling, and cross-dimension interaction detection. Evaluating five frontier LLMs across 7,225 dialogues (85 cases x 17 configurations x 5 models), we find a fundamental asymmetry: information pollution (fabricating symptoms) produces 1.7-3.4x larger accuracy drops than information deficit (withholding information), and fabricating is the only configuration achieving statistical significance across all five models (McNemar p < 0.05). Among six dimension combinations, fabricating is the sole driver of super-additive interaction: all three fabricating-involving pairs produce O/E ratios of 0.70-0.81 (35-44% of eligible cases fail under the combination despite succeeding under each dimension alone), while all non-fabricating pairs show purely additive effects (O/E ~ 1.0). Inquiry strategy moderates deficit but not pollution: exhaustive questioning recovers withheld information, but cannot compensate for fabricated inputs. Models exhibit distinct vulnerability profiles, with worst-case drops ranging from 38.8 to 54.1 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。