arXiv:2512.22974cs.CV2025-12被引 1

RealCamo通过布局控制与图文引导,生成更真实、语义一致的伪装图像。

RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance

  • 基于外绘框架,引入布局控制调节全局结构。
  • 结合细粒度文本描述与纹理检索,提升背景视觉真实性。
  • 提出分布差异度量,定量评估伪装效果,适合数据增强研究者。

伪装图像生成(CIG)近年来成为获取高质量伪装目标检测(COD)训练数据的有效替代方案。然而,现有CIG方法仍与真实伪装图像存在显著差距:生成图像或因视觉相似性弱导致伪装不足,或出现语义不一致的杂乱背景。为此,我们提出RealCamo,一种基于外绘的可控真实伪装图像生成框架。RealCamo显式引入额外布局控制以调节全局图像结构,从而提升前景物体与生成背景间的语义一致性。此外,我们构建多模态文本-视觉条件,结合统一的细粒度文本任务描述与面向纹理的背景检索,协同引导生成过程以增强视觉保真度与真实感。为定量评估伪装质量,我们进一步提出背景-前景分布差异度量,用于衡量生成图像中伪装的有效性。大量实验与可视化结果验证了所提框架的有效性。

原文摘要 · Abstract (English)

Camouflaged image generation (CIG) has recently emerged as an efficient alternative for acquiring high-quality training data for camouflaged object detection (COD). However, existing CIG methods still suffer from a substantial gap to real camouflaged imagery: generated images either lack sufficient camouflage due to weak visual similarity, or exhibit cluttered backgrounds that are semantically inconsistent with foreground targets. To address these limitations, we propose RealCamo, a novel out-painting-based framework for controllable realistic camouflaged image generation. RealCamo explicitly introduces additional layout controls to regulate global image structure, thereby improving semantic coherence between foreground objects and generated backgrounds. Moreover, we construct a multimodal textual-visual condition by combining a unified fine-grained textual task description with texture-oriented background retrieval, which jointly guides the generation process to enhance visual fidelity and realism. To quantitatively assess camouflage quality, we further introduce a background-foreground distribution divergence metric that measures the effectiveness of camouflage in generated images. Extensive experiments and visualizations demonstrate the effectiveness of our proposed framework.

图像生成伪装检测多模态布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。