提出量化动物伪装程度的新方法,让机器判断伪装效果更接近人类直觉。
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability

- 将伪装评估转为视觉定位问题,用类别信息生成检测候选框。
- 在2390次对比中与人眼判断一致率达78.82%,比现有方法高25%。
- 可作为训练生成模型的偏好信号,适合用于生成自然伪装图像。
动物在环境中无缝融合时被视为有效伪装,但目前缺乏标准化的定量评估方法。本文提出SeamCam(无缝伪装)度量方法,将伪装评价建模为视觉定位任务:当已知物种类别时仍难以发现的动物即为强伪装。SeamCam通过生成类别条件下的检测建议框、提取分割掩码,并找出使联合区域与真实掩码交并比最高的子集,其得分定义为1减去该最大可恢复定位信号,分数越高表示伪装越强。在包含94名参与者和2390次比较的人类双选择实验中,SeamCam与人类对伪装难度的判断一致性达78.82%,优于当前最优方法约25%。进一步展示其可用于直接偏好优化(DPO)微调基于扩散的修复模型,实现专为伪装生成设计的低成本训练。为此还构建了高质量数据集CamFG-1.5k,包含1521张高分辨率图像,动物在伪装前完全可见,避免了现有数据集中遮挡伪影的影响,支持无偏评估。
原文摘要 · Abstract (English)
Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this seamlessness exists. We address this gap by framing camouflage evaluation as a visual localization problem: a well-camouflaged animal is one that remains difficult to detect even when its category is known. We introduce SeamCam (Seamless Camouflage), a metric that quantifies how detectable an animal is from the available visual evidence. Given an image and a target species, SeamCam generates category-conditioned detection proposals, extracts segmentation masks, and identifies the subset whose collective union yields the highest IoU with the ground-truth mask. The SeamCam score is one minus this maximum recoverable localization signal, where a higher score indicates stronger camouflage (i.e., lower detectability). In a human two-alternative forced-choice study with 94 participants and 2,390 comparisons, SeamCam achieves 78.82% agreement with human camouflage difficulty judgments, outperforming state-of-the-art by about 25%. We then demonstrate SeamCam's utility as a preference signal for Direct Preference Optimization (DPO) to fine-tune a diffusion-based inpainting model for camouflage generation. This offers an affordable training approach with an objective explicitly suited for camouflage generation, unlike typical diffusion models. To support rigorous benchmarking, we further introduce CamFG-1.5k, a curated dataset of 1,521 high-resolution images in which animals are fully visible prior to camouflage generation, enabling unbiased evaluation by controlling for occlusion artifacts present in existing datasets. https://7amin.github.io/SeamCam/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。