用生成模型自动生成高质量伪装图像密集标注数据
GenCAMO: Scene-Graph Contextual Decoupling for Environment-aware and Mask-free Camouflage Image-Dense Annotation Generation
- 基于场景图上下文解耦,实现环境感知的无掩码生成
- 生成数据使复杂伪装场景的密集预测性能显著提升
- 适合需要大量真实标注数据的视觉理解研究者
伪装密集预测(CDP),尤其是RGB-D伪装目标检测和开放词汇伪装目标分割,在提升复杂伪装场景的理解与推理能力方面至关重要。然而,由于数据采集和标注成本高昂,高质量且大规模带有密集标注的伪装数据集仍然稀缺。为解决这一挑战,我们探索利用生成模型合成逼真的伪装图像密集数据,以训练具备细粒度表征、先验知识和辅助推理能力的CDP模型。具体贡献有三:(i) 提出GenCAMO-DB,一个包含深度图、场景图、属性描述和文本提示等多模态标注的大规模伪装数据集;(ii) 提出GenCAMO,一种环境感知且无需掩码的生成框架,可生成高保真伪装图像密集标注;(iii) 多模态实验表明,GenCAMO通过提供高质量合成数据,显著提升了复杂伪装场景下的密集预测性能。代码与数据集将在论文接受后公开。
原文摘要 · Abstract (English)
Conceal dense prediction (CDP), especially RGB-D camouflage object detection and open-vocabulary camouflage object segmentation, plays a crucial role in advancing the understanding and reasoning of complex camouflage scenes. However, high-quality and large-scale camouflage datasets with dense annotation remain scarce due to expensive data collection and labeling costs. To address this challenge, we explore leveraging generative models to synthesize realistic camouflage image-dense data for training CDP models with fine-grained representations, prior knowledge, and auxiliary reasoning. Concretely, our contributions are threefold: (i) we introduce GenCAMO-DB, a large-scale camouflage dataset with multi-modal annotations, including depth maps, scene graphs, attribute descriptions, and text prompts; (ii) we present GenCAMO, an environment-aware and mask-free generative framework that produces high-fidelity camouflage image-dense annotations; (iii) extensive experiments across multiple modalities demonstrate that GenCAMO significantly improves dense prediction performance on complex camouflage scenes by providing high-quality synthetic data. The code and datasets will be released after paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。