arXiv:2506.06748cs.CV2025-06
用视觉与深度信息融合,提升第一人称视频目标分割精度
THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
- 结合SAM2预训练与深度几何线索,统一建模视频对象
- 在VISOR测试集上实现90.1%的J&F分数,性能领先
- 适合关注第一人称视频理解与长时跟踪的应用场景
本文描述了我们在第一人称视频目标分割任务中的方法。通过将SAM2的大规模视觉预训练与基于深度的几何线索相结合,我们的方法能够有效处理复杂场景和长期追踪问题。在统一框架中融合多源信号,显著提升了分割性能。在VISOR测试集上,本方法达到90.1%的J&F得分,展现了优异的鲁棒性与准确性。
原文摘要 · Abstract (English)
In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to handle complex scenes and long-term tracking. By integrating these signals in a unified framework, we achieve strong segmentation performance. On the VISOR test set, our method reaches a J&F score of 90.1%.
视频分割第一人称视频深度线索
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。