通过神经元消融发现,医疗大模型角色扮演仅改语言风格,未提升推理能力。
Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
- 用神经元消融与表征分析检验角色提示对推理路径的影响
- 三组医疗问答数据集验证:角色提示未引发不同认知路径
- 适合关注医疗AI认知真实性与角色扮演局限的研究者
大型语言模型在医疗决策支持系统中应用广泛,尤其在医疗问答和角色扮演模拟中。常见的基于提示的角色扮演(PBRP)要求模型扮演不同临床角色(如医学生、住院医师、主治医师),以模拟多样化专业行为。然而,角色提示对模型推理能力的影响尚不明确。本研究提出RP-Neuron-Activated Evaluation框架(RPNA),评估角色提示是否引发特定于角色的认知过程,或仅改变语言风格。我们在三个医疗QA数据集上测试该框架,结合神经元消融与表征分析技术,评估推理路径变化。结果表明,角色提示并未显著提升模型的医学推理能力,主要影响表面语言特征,未发现不同临床角色间存在差异化的推理路径或认知区分。尽管存在表面风格变化,模型的核心决策机制在各类角色中保持一致,说明当前PBRP方法无法复现真实医疗实践中的认知复杂性。这揭示了角色扮演在医疗AI中的局限性,强调需发展能模拟真实认知过程而非语言模仿的模型。相关代码已开源:https://github.com/IAAR-Shanghai/RolePlay_LLMDoctor
原文摘要 · Abstract (English)
Large language models (LLMs) have gained significant traction in medical decision support systems, particularly in the context of medical question answering and role-playing simulations. A common practice, Prompt-Based Role Playing (PBRP), instructs models to adopt different clinical roles (e.g., medical students, residents, attending physicians) to simulate varied professional behaviors. However, the impact of such role prompts on model reasoning capabilities remains unclear. This study introduces the RP-Neuron-Activated Evaluation Framework(RPNA) to evaluate whether role prompts induce distinct, role-specific cognitive processes in LLMs or merely modify linguistic style. We test this framework on three medical QA datasets, employing neuron ablation and representation analysis techniques to assess changes in reasoning pathways. Our results demonstrate that role prompts do not significantly enhance the medical reasoning abilities of LLMs. Instead, they primarily affect surface-level linguistic features, with no evidence of distinct reasoning pathways or cognitive differentiation across clinical roles. Despite superficial stylistic changes, the core decision-making mechanisms of LLMs remain uniform across roles, indicating that current PBRP methods fail to replicate the cognitive complexity found in real-world medical practice. This highlights the limitations of role-playing in medical AI and emphasizes the need for models that simulate genuine cognitive processes rather than linguistic imitation.We have released the related code in the following repository:https: //github.com/IAAR-Shanghai/RolePlay_LLMDoctor
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。