AI生成的皮肤科治疗方案被人类和更高级AI评价时,得分完全相反,暴露了临床直觉与算法逻辑的根本差异。
Divergent Realities: A Comparative Analysis of Human Expert vs. Artificial Intelligence Based Generation and Evaluation of Treatment Plans in Dermatology
- 对比人类专家与两类AI生成的治疗方案,用相同标准评分
- 人类专家偏爱人类方案(均分7.62对7.16),而高级AI更青睐AI方案(均分7.75对6.79)
- 推理型AI o3在人类眼中排第11,在AI评委中却排名第一,凸显评价体系差异
背景:随着AI从诊断扩展到治疗规划,评估其生成方案成为关键挑战,尤其在新型推理模型出现后。本研究对比了10位皮肤科专家、通用型AI(GPT-4o)和推理型AI(o3)针对5个复杂病例生成的治疗方案,并由人类同行与更先进的AI裁判(Gemini 2.5 Pro)进行双阶段评分。方法:所有方案匿名化并标准化后,先由10名专家打分,再由同一套评分标准下的高级AI裁判打分。结果:显著的‘评价者效应’显现——人类专家对同类方案评分更高(均分7.62 vs. 7.16;p=0.0313),GPT-4o排第6(均分7.38),o3排第11(均分6.97)。而AI裁判则出现完全反转:对AI方案评分更高(均分7.75 vs. 6.79;p=0.0313),将o3列为第1(均分8.20),GPT-4o第2,所有人类专家均被置于较低位置。结论:临床方案的质量感知高度依赖于评价者类型。先进推理型AI虽被人类低估,却被高级AI评为最优,揭示经验性临床直觉与数据驱动算法逻辑之间的深层鸿沟。这一悖论对AI融入医疗构成重大挑战,表明未来需构建可解释、协同的人机系统以弥合推理差异,增强临床决策支持。
原文摘要 · Abstract (English)
Background: Evaluating AI-generated treatment plans is a key challenge as AI expands beyond diagnostics, especially with new reasoning models. This study compares plans from human experts and two AI models (a generalist and a reasoner), assessed by both human peers and a superior AI judge. Methods: Ten dermatologists, a generalist AI (GPT-4o), and a reasoning AI (o3) generated treatment plans for five complex dermatology cases. The anonymized, normalized plans were scored in two phases: 1) by the ten human experts, and 2) by a superior AI judge (Gemini 2.5 Pro) using an identical rubric. Results: A profound 'evaluator effect' was observed. Human experts scored peer-generated plans significantly higher than AI plans (mean 7.62 vs. 7.16; p=0.0313), ranking GPT-4o 6th (mean 7.38) and the reasoning model, o3, 11th (mean 6.97). Conversely, the AI judge produced a complete inversion, scoring AI plans significantly higher than human plans (mean 7.75 vs. 6.79; p=0.0313). It ranked o3 1st (mean 8.20) and GPT-4o 2nd, placing all human experts lower. Conclusions: The perceived quality of a clinical plan is fundamentally dependent on the evaluator's nature. An advanced reasoning AI, ranked poorly by human experts, was judged as superior by a sophisticated AI, revealing a deep gap between experience-based clinical heuristics and data-driven algorithmic logic. This paradox presents a critical challenge for AI integration, suggesting the future requires synergistic, explainable human-AI systems that bridge this reasoning gap to augment clinical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。