arXiv:2509.15596cs.CV2025-09中稿 · NeurIPS被引 3

构建眼科手术认知评估基准,测试模型看图、懂知识、做判断的能力

EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery

  • 基于210万+视觉问答,覆盖上千细粒度属性,评估多视角感知能力
  • 医疗知识图谱含2.5万+三元组,支持医学知识理解与推理任务
  • 首次系统评测大模型在眼科手术中的认知表现,适合临床视觉智能研究者

多模态大模型虽表现出色,但在高风险、专业性强的手术场景中表现仍不明确。为此,我们构建了面向眼科手术分析的大型基准测试集EyePCR,基于结构化临床知识,评估模型在感知、理解与推理三个层面的认知能力。EyePCR包含超过210万条视觉问答,涵盖1048个细粒度属性用于多视角感知,拥有超过2.5万个三元组的医学知识图谱支持知识理解,并设计了四项临床相关的推理任务。丰富的标注支持对医生如何结合视觉线索与领域知识进行决策的深度分析,显著提升模型的认知能力。特别地,经领域适配的EyePCR-MLLM(基于Qwen2.5-VL-7B)在感知类多项选择题中达到最优准确率,在理解和推理任务上优于开源模型,接近GPT-4.1等商业模型表现。EyePCR揭示了现有多模态大模型在手术认知上的局限性,为外科视频理解模型的评估与临床可靠性提升奠定基础。

原文摘要 · Abstract (English)

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap, we develop \textbf{EyePCR}, a large-scale benchmark for ophthalmic surgery analysis, grounded in structured clinical knowledge to evaluate cognition across \textit{Perception}, \textit{Comprehension} and \textit{Reasoning}. EyePCR offers a richly annotated corpus with more than 210k VQAs, which cover 1048 fine-grained attributes for multi-view perception, medical knowledge graph of more than 25k triplets for comprehension, and four clinically grounded reasoning tasks. The rich annotations facilitate in-depth cognitive analysis, simulating how surgeons perceive visual cues and combine them with domain knowledge to make decisions, thus greatly improving models' cognitive ability. In particular, \textbf{EyePCR-MLLM}, a domain-adapted variant of Qwen2.5-VL-7B, achieves the highest accuracy on MCQs for \textit{Perception} among compared models and outperforms open-source models in \textit{Comprehension} and \textit{Reasoning}, rivalling commercial models like GPT-4.1. EyePCR reveals the limitations of existing MLLMs in surgical cognition and lays the foundation for benchmarking and enhancing clinical reliability of surgical video understanding models.

眼科手术多模态大模型认知评估视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。