评测大模型在动态牙科诊疗中的表现,发现其决策可靠性不足。
Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios
- 构建牙科临床模拟评估基准,测试静态任务到多轮对话的全流程表现。
- 模型在动态对话中表现骤降,主要瓶颈是信息收集与状态追踪能力弱。
- 外部知识增强对动态任务效果有限,需领域适配预训练才能提升安全决策。
大型语言模型(LLMs)从被动知识检索者向自主临床代理的转变,要求评估方式从静态准确率转向动态行为可靠性。为探索牙科领域这一边界——该领域高质量AI建议能促进患者参与决策——我们提出标准化临床管理与性能评估(SCMPE)基准,全面评估模型在知识导向任务(静态客观任务)到工作流模拟(多轮模拟患者互动)的表现。分析显示,尽管模型在静态任务中表现优异,但在动态临床对话中性能急剧下降,表明主要瓶颈并非知识保留,而是主动信息获取和动态状态追踪的挑战。通过“指南遵循度”与“决策质量”的映射发现,通用模型普遍存在“高疗效、低安全性”风险。此外,我们量化了检索增强生成(RAG)的影响:虽然在静态任务中可减少幻觉,但在动态工作流中效果有限且不一致,有时甚至导致性能退化。这表明,仅靠外部知识无法弥补推理差距,必须结合领域自适应预训练。本研究实证刻画了牙科LLMs的能力边界,为弥合标准知识与安全自主临床实践之间的鸿沟提供了路线图。
原文摘要 · Abstract (English)
The transition of Large Language Models (LLMs) from passive knowledge retrievers to autonomous clinical agents demands a shift in evaluation-from static accuracy to dynamic behavioral reliability. To explore this boundary in dentistry, a domain where high-quality AI advice uniquely empowers patient-participatory decision-making, we present the Standardized Clinical Management & Performance Evaluation (SCMPE) benchmark, which comprehensively assesses performance from knowledge-oriented evaluations (static objective tasks) to workflow-based simulations (multi-turn simulated patient interactions). Our analysis reveals that while models demonstrate high proficiency in static objective tasks, their performance precipitates in dynamic clinical dialogues, identifying that the primary bottleneck lies not in knowledge retention, but in the critical challenges of active information gathering and dynamic state tracking. Mapping "Guideline Adherence" versus "Decision Quality" reveals a prevalent "High Efficacy, Low Safety" risk in general models. Furthermore, we quantify the impact of Retrieval-Augmented Generation (RAG). While RAG mitigates hallucinations in static tasks, its efficacy in dynamic workflows is limited and heterogeneous, sometimes causing degradation. This underscores that external knowledge alone cannot bridge the reasoning gap without domain-adaptive pre-training. This study empirically charts the capability boundaries of dental LLMs, providing a roadmap for bridging the gap between standardized knowledge and safe, autonomous clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。