arXiv:2512.11544cs.AI2025-12被引 1

首次发现大模型处理临床文本时会像脂肪肝一样‘功能失常’。

AI-MASLD Metabolic Dysfunction and Information Steatosis of Large Language Models in Unstructured Clinical Narratives

  • 用真实临床语料测试四大主流大模型,模拟医生问诊场景。
  • 所有模型在高噪声下性能骤降,GPT-4o误判肺栓塞风险。
  • 提出‘AI-MASLD’概念,警示医疗应用需人工监督。

本研究旨在模拟真实临床场景,系统评估大语言模型(LLMs)从充满噪声和冗余的患者主诉中提取核心医学信息的能力,并验证其是否表现出类似代谢功能障碍相关脂肪肝病(MASLD)的功能衰退。采用基于标准化医学探针的横断面分析设计,选取GPT-4o、Gemini 2.5、DeepSeek 3.1和Qwen3-Max四款主流模型作为研究对象。构建包含二十个医学探针、覆盖五个核心维度的评估体系,以模拟真实的临床沟通环境。所有探针均设有临床专家定义的标准答案,并由两名独立医师采用双盲逆向评分法进行评估。结果表明,所有测试模型均存在不同程度的功能缺陷,其中Qwen3-Max表现最佳,Gemini 2.5最差。在极端噪声条件下,多数模型出现功能崩溃。值得注意的是,GPT-4o在深静脉血栓(DVT)引发肺栓塞(PE)的风险评估中出现严重误判。本研究首次实证证明大模型在处理临床信息时具有类似代谢功能障碍的特征,提出‘人工智能代谢功能障碍相关脂肪肝病(AI-MASLD)’这一创新概念。研究为人工智能在医疗领域的应用敲响安全警钟,强调当前大模型仍需在人类专家监督下作为辅助工具使用,其理论知识与实际临床应用之间仍存在显著差距。

原文摘要 · Abstract (English)

This study aims to simulate real-world clinical scenarios to systematically evaluate the ability of Large Language Models (LLMs) to extract core medical information from patient chief complaints laden with noise and redundancy, and to verify whether they exhibit a functional decline analogous to Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD). We employed a cross-sectional analysis design based on standardized medical probes, selecting four mainstream LLMs as research subjects: GPT-4o, Gemini 2.5, DeepSeek 3.1, and Qwen3-Max. An evaluation system comprising twenty medical probes across five core dimensions was used to simulate a genuine clinical communication environment. All probes had gold-standard answers defined by clinical experts and were assessed via a double-blind, inverse rating scale by two independent clinicians. The results show that all tested models exhibited functional defects to varying degrees, with Qwen3-Max demonstrating the best overall performance and Gemini 2.5 the worst. Under conditions of extreme noise, most models experienced a functional collapse. Notably, GPT-4o made a severe misjudgment in the risk assessment for pulmonary embolism (PE) secondary to deep vein thrombosis (DVT). This research is the first to empirically confirm that LLMs exhibit features resembling metabolic dysfunction when processing clinical information, proposing the innovative concept of "AI-Metabolic Dysfunction-Associated Steatotic Liver Disease (AI-MASLD)". These findings offer a crucial safety warning for the application of Artificial Intelligence (AI) in healthcare, emphasizing that current LLMs must be used as auxiliary tools under human expert supervision, as there remains a significant gap between their theoretical knowledge and practical clinical application.

大模型评测医疗AIAI安全临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。