arXiv:2511.20218cs.CV2025-11AAAI被引 2

用文本控制生成更自然的伪装图像,让物体与环境逻辑融合。

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

  • 通过视觉语言模型生成高质量文本提示,指导图像生成
  • 引入轻量控制器和高频特征模块,提升伪装真实度
  • 适合需要高保真伪装图像的军事或安全研究场景

伪装图像生成(CIG)是新兴研究方向,旨在合成使物体与周围环境和谐融合、视觉一致性高的图像。现有方法或直接将物体嵌入特定背景,或通过前景引导扩散模型扩展环境,但常因忽略伪装物与背景间的逻辑关系而生成不自然结果。为此,本文提出可控文本引导伪装图像生成方法CT-CIG,利用大视觉语言模型设计伪装揭示对话机制(CRDM),为现有伪装数据集添加高质量文本提示。构建的图像-提示对用于微调Stable Diffusion,引入轻量控制器以指导伪装物体的位置与形状,增强场景契合度。此外,设计频率交互精炼模块(FIRM)捕捉高频纹理特征,促进复杂伪装模式的学习。大量实验表明,生成提示具有语义一致性,且CT-CIG可生成逼真的伪装图像,经CLIPScore评估和伪装效果测试验证。

原文摘要 · Abstract (English)

Camouflage Images Generation (CIG) is an emerging research area that focuses on synthesizing images in which objects are harmoniously blended and exhibit high visual consistency with their surroundings. Existing methods perform CIG by either fusing objects into specific backgrounds or outpainting the surroundings via foreground object-guided diffusion. However, they often fail to obtain natural results because they overlook the logical relationship between camouflaged objects and background environments. To address this issue, we propose CT-CIG, a Controllable Text-guided Camouflage Images Generation method that produces realistic and logically plausible camouflage images. Leveraging Large Visual Language Models (VLM), we design a Camouflage-Revealing Dialogue Mechanism (CRDM) to annotate existing camouflage datasets with high-quality text prompts. Subsequently, the constructed image-prompt pairs are utilized to finetune Stable Diffusion, incorporating a lightweight controller to guide the location and shape of camouflaged objects for enhanced camouflage scene fitness. Moreover, we design a Frequency Interaction Refinement Module (FIRM) to capture high-frequency texture features, facilitating the learning of complex camouflage patterns. Extensive experiments, including CLIPScore evaluation and camouflage effectiveness assessment, demonstrate the semantic alignment of our generated text prompts and CT-CIG's ability to produce photorealistic camouflage images.

图像生成扩散模型伪装技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。