构建了大规模RGB-D视频显著目标检测数据集RDVSv2,支持更精准的多模态视频理解。
RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

- 基于立体视频生成带深度图的标注数据,融合眼动追踪指导的掩码。
- 包含249个视频、29,077帧,规模远超现有数据集,场景更复杂多样。
- 基于SAM2设计高效微调方案,显著提升多模态视频显著性检测性能。
我们提出了RDVSv2,一个大规模的RGB-D视频显著目标检测(RGB-D VSOD)基准,包含密集的帧级标注。现有该领域数据集通常规模有限、标注质量不高,且依赖的深度线索几何一致性较弱。为解决这些问题,RDVSv2从公开可获取的立体视频中构建,包含249个视频序列,共29,077帧标注数据,提供由立体视频推导出的深度图,以及结合眼动追踪引导的逐帧显著目标掩码。相比现有数据集,RDVSv2在规模上显著更大,覆盖更多样且更具挑战性的场景。此外,我们基于Segment Anything Model 2(SAM2)建立了一个强大的基准方法,采用参数高效微调(PEFT)策略,使SAM2编码器能联合编码RGB、深度和光流信息。大量实验表明,RDVSv2对现有RGB-D VSOD方法构成显著挑战;同时,所提基准在RDVSv2及现有数据集上均达到领先性能。我们希望该数据集与基线能为未来RGB-D VSOD及相关多模态视频理解研究提供有力支持。数据与代码将开源于https://github.com/ltynick/RDVSv2。
原文摘要 · Abstract (English)
We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets in this emerging field are often limited in scale and annotation quality, while also relying on less geometry-consistent depth cues. To address these limitations, RDVSv2 is built from publicly accessible stereoscopic online videos and contains 249 video sequences with 29,077 annotated frames. It includes depth maps derived from stereoscopic videos, together with frame-wise salient object masks annotated with eye-tracking guidance. Compared with existing datasets, RDVSv2 is much larger in scale and covers more diverse and challenging scenarios. In addition, we establish a strong baseline for RGB-D VSOD based on Segment Anything Model 2 (SAM2). Specifically, we employ a parameter-efficient fine-tuning (PEFT) strategy to adapt the SAM2 encoder to jointly encode RGB, depth, and optical flow cues. Extensive experiments show that RDVSv2 is substantially more challenging for existing RGB-D VSOD methods. Meanwhile, the proposed baseline achieves state-of-the-art results on RDVSv2 and existing RGB-D VSOD benchmarks. We hope that RDVSv2 and the provided baseline will serve as useful resources for future research on RGB-D VSOD and related multi-modal video understanding tasks. Our dataset and code will be available at https://github.com/ltynick/RDVSv2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。