arXiv:2507.00339cs.CVcs.AI2025-07

首个支持多视角的无模态内容重建数据集,助力复杂场景下物体完整表征。

Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video

  • 构建多摄像头视频数据集,实现跨视角物体身份一致追踪
  • 标注超580万实例,首次提供无模态内容真实标签
  • 适合研究视觉推理、3D重建与多视角感知的学者使用

无模态分割与无模态内容补全需依赖物体先验来推断复杂场景中被遮挡的掩码和特征。现有数据缺乏多摄像头共享场景视角的上下文维度。本文提出MOVi-MC-AC:包含多对象、多相机与无模态内容的大型数据集,是迄今最大的无模态分割数据集,也是首个提供无模态内容真实标签的数据集。在合成多摄像头视频中模拟通用家用物品的杂乱场景,为物体检测、跟踪与分割研究提供新基准。该数据集通过为每个物体分配一致的实例ID,在不同帧与摄像头间保持身份一致性,且每台相机具有独特运动模式与特征。无模态内容任务要求模型预测被遮挡物体的外观。以往方法依赖缓慢的拼贴生成伪标签,无法反映真实遮挡关系。本数据集提供约580万实例的标注,开创性地建立无模态内容真实标签体系,相关资源已公开于https://huggingface.co/datasets/Amar-S/MOVi-MC-AC。

原文摘要 · Abstract (English)

Amodal segmentation and amodal content completion require using object priors to estimate occluded masks and features of objects in complex scenes. Until now, no data has provided an additional dimension for object context: the possibility of multiple cameras sharing a view of a scene. We introduce MOVi-MC-AC: Multiple Object Video with Multi-Cameras and Amodal Content, the largest amodal segmentation and first amodal content dataset to date. Cluttered scenes of generic household objects are simulated in multi-camera video. MOVi-MC-AC contributes to the growing literature of object detection, tracking, and segmentation by including two new contributions to the deep learning for computer vision world. Multiple Camera (MC) settings where objects can be identified and tracked between various unique camera perspectives are rare in both synthetic and real-world video. We introduce a new complexity to synthetic video by providing consistent object ids for detections and segmentations between both frames and multiple cameras each with unique features and motion patterns on a single scene. Amodal Content (AC) is a reconstructive task in which models predict the appearance of target objects through occlusions. In the amodal segmentation literature, some datasets have been released with amodal detection, tracking, and segmentation labels. While other methods rely on slow cut-and-paste schemes to generate amodal content pseudo-labels, they do not account for natural occlusions present in the modal masks. MOVi-MC-AC provides labels for ~5.8 million object instances, setting a new maximum in the amodal dataset literature, along with being the first to provide ground-truth amodal content. The full dataset is available at https://huggingface.co/datasets/Amar-S/MOVi-MC-AC ,

无模态分割多视角数据集内容补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。