模仿人类逐个查看的顺序,提升手术器械密集场景下的计数准确率。
Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting
- 通过构建视觉链路模拟人眼逐个扫描过程,避免传统检测的无序性。
- 在1464张高密度图像上,显著超越现有方法和大模型的计数表现。
- 适用于需要精准器械清点的智能手术室系统,尤其适合密集场景。
手术室中精确计数手术器械是保障患者安全的关键前提。尽管大视觉语言模型和代理式AI取得进展,但在器械紧密聚集的密集场景下,准确计数仍具挑战性。为此,本文提出Chain-of-Look,一种模拟人类有序观察过程的视觉推理框架,通过强制构建结构化视觉链,替代传统的无序目标检测。该视觉链引导模型沿连贯的空间轨迹进行计数,提升复杂场景下的准确性。为进一步增强视觉链的物理合理性,引入邻近损失函数,显式建模密集器械间的空间约束。同时构建SurgCount-HD数据集,包含1,464张高密度手术器械图像。大量实验表明,本方法在密集器械计数任务中优于当前最优方法(如CountGD、REC)以及多模态大模型(如Qwen、ChatGPT)。
原文摘要 · Abstract (English)
Accurate counting of surgical instruments in Operating Rooms (OR) is a critical prerequisite for ensuring patient safety during surgery. Despite recent progress of large visual-language models and agentic AI, accurately counting such instruments remains highly challenging, particularly in dense scenarios where instruments are tightly clustered. To address this problem, we introduce Chain-of-Look, a novel visual reasoning framework that mimics the sequential human counting process by enforcing a structured visual chain, rather than relying on classic object detection which is unordered. This visual chain guides the model to count along a coherent spatial trajectory, improving accuracy in complex scenes. To further enforce the physical plausibility of the visual chain, we introduce the neighboring loss function, which explicitly models the spatial constraints inherent to densely packed surgical instruments. We also present SurgCount-HD, a new dataset comprising 1,464 high-density surgical instrument images. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches for counting (e.g., CountGD, REC) as well as Multimodality Large Language Models (e.g., Qwen, ChatGPT) in the challenging task of dense surgical instrument counting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。