arXiv:2601.06750cs.CVcs.AI2026-01ACL被引 1

首个用医生视线评估医疗多模态大模型意图理解能力的基准

Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

  • 以医生注视点作为认知光标,评估模型在手术等场景中的意图理解
  • 发现现有模型因依赖全局特征而易编造观察结果、盲目服从指令
  • 引入陷阱问答机制,检测幻觉和盲从,提升临床可靠性测试精度

医疗多模态大语言模型(Med-MLLMs)在真实场景部署中需具备以患者为中心的临床意图理解能力,但现有基准无法有效评估这一关键能力。为此,我们提出 MedGaze-Bench,首个利用医生注视点作为认知光标来评估手术、急诊模拟和诊断解读中意图理解的基准。该基准解决三大挑战:解剖结构视觉同质性、临床流程严格的时间因果依赖、以及隐式安全规范遵循。我们构建三维临床意图框架,评估:(1) 空间意图:在视觉噪声中精准识别目标;(2) 时间意图:通过回溯与前瞻推理推断因果逻辑;(3) 标准意图:通过安全检查验证规程合规性。除准确率外,引入陷阱问答(Trap QA)机制,通过惩罚幻觉和认知逢迎来压力测试临床可靠性。实验表明,当前 MLLMs 因过度依赖全局特征,在以我为中心意图理解上表现不佳,常生成虚构观察并盲目接受无效指令。

原文摘要 · Abstract (English)

Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.

医疗AI多模态意图理解基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。