arXiv:2507.18517cs.CV2025-07被引 1

用眼球注视引导大模型实现杂乱场景中物体分割,提升假肢视觉控制精度。

Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs

  • 基于眼动数据生成提示,指导SAM模型在复杂场景中分割物体。
  • 在Grasping-in-the-Wild数据集上IoU提升0.51点,实测效果显著。
  • 适用于上肢神经假肢的视觉辅助系统,无需额外微调即可部署。

本文研究利用基础模型进行语义物体分割的问题。我们探讨了在大量且多样化的物体上训练的基础模型,是否能在不针对特定图像进行微调的情况下,对包含日常物品的高杂乱视觉场景实现物体分割。该“野外”情境源于视觉引导上肢神经假肢的应用需求。我们提出一种基于眼球注视生成提示的方法,以指导分割任意模型(SAM)在分割任务中的表现,并在第一人称视角数据上对其进行微调。评估结果表明,本方法在公开于RoboFlow平台的Grasping-in-the-Wild数据集上的真实世界挑战性数据中,将分割质量指标IoU提升了最高0.51点。

原文摘要 · Abstract (English)

In this work, we address the problem of semantic object segmentation using foundation models. We investigate whether foundation models, trained on a large number and variety of objects, can perform object segmentation without fine-tuning on specific images containing everyday objects, but in highly cluttered visual scenes. The ''in the wild'' context is driven by the target application of vision guided upper limb neuroprostheses. We propose a method for generating prompts based on gaze fixations to guide the Segment Anything Model (SAM) in our segmentation scenario, and fine-tune it on egocentric visual data. Evaluation results of our approach show an improvement of the IoU segmentation quality metric by up to 0.51 points on real-world challenging data of Grasping-in-the-Wild corpus which is made available on the RoboFlow Platform (https://universe.roboflow.com/iwrist/grasping-in-the-wild)

物体分割神经假肢基础模型眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。