用合成数据解决手术室视觉-语义知识冲突问题,提升AI安全识别能力。
OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
- 基于真实手术室数据生成2.8万张高保真合成图像,构建专属基准
- 实测主流多模态大模型在手术风险识别上仍存在显著可靠性差距
- 微调后模型可有效缓解知识冲突,跨视角泛化能力显著提升
自动化识别手术安全风险对改善患者预后至关重要;然而,多模态大语言模型(MLLMs)常出现视觉-语义知识冲突(VS-KC),即虽具备安全知识但无法在视觉检查中激活。研究手术室中的此类对齐差距受限于真实数据稀缺与隐私约束。为此,我们提出OR-VSKC基准,用于研究严格监管环境下的手术风险感知与VS-KC。该基准基于我们的协议到像素生成框架构建,包含28,190张符合权威安全标准的高保真合成图像,并附有713张专家标注的挑战性子集,经多位专家验证。整个数据集源自4D-OR和CAMMA-MVOR真实手术室数据,其中4D-OR部分作为主基准核心,CAMMA-MVOR部分保留用于外部验证与跨数据集泛化分析。对先进MLLMs的评估显示,即使在顶级通用模型中也存在显著可靠性差距。实验表明,在OR-VSKC上微调可有效缓解VS-KC,并实现对未见摄像头视角的稳健泛化。代码与数据已开源,支持医疗高危场景的可复现研究。源码地址:https://github.com/zgg2577/VS-KC。
原文摘要 · Abstract (English)
Automated identification of surgical safety risks is critical for improving patient outcomes; however, Multimodal Large Language Models (MLLMs) frequently suffer from Visual-Semantic Knowledge Conflicts (VS-KC), a phenomenon where models possess safety knowledge but fail to activate it during visual inspection. Investigating this alignment gap in operating rooms (ORs) is impeded by a critical bottleneck: the scarcity and privacy constraints of real-world OR data depicting safety violations. To address this, we introduce OR-VSKC, a benchmark for studying VS-KC and surgical risk perception in strictly regulated OR environments. Constructed via our Protocol-to-Pixel Generative Framework, OR-VSKC comprises 28,190 high-fidelity synthetic images grounded in authoritative safety standards, complemented by a 713-image expert-authored challenge subset validated by multiple experts. The full benchmark is built from authentic OR contexts drawn from the 4D-OR and CAMMA-MVOR datasets, where the 4D-OR-based portion serves as the primary benchmark core and the CAMMA-MVOR-based portion is reserved for external validation and cross-dataset generalization analysis. Evaluations of state-of-the-art MLLMs reveal substantial reliability gaps even in advanced generalist models. Furthermore, experiments show that fine-tuning on OR-VSKC effectively mitigates VS-KC and enables robust generalization to unseen camera viewpoints. We open-source the code and dataset to support reproducible research in safety-critical medical environments. The source code is available at https://github.com/zgg2577/VS-KC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。