让AI像人一样看穿伪装物体,通过逐步聚焦提升识别能力
Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
- 设计视觉聚焦强化框架,引导模型分步推理定位隐蔽物体
- 在伪装物体识别与检测任务中显著优于监督微调基线
- 适合研究视觉认知、模型可解释性及高级视觉推理的学者
当前多模态模型在识别与背景融合的隐蔽物体时,与人类视觉系统存在明显偏差。我们发现这些模型无法区分隐藏物体,难以模拟人类利用前景-背景相似性进行视觉分析的认知过程。为此,我们构建了模仿人类伪装感知的视觉系统,通过迭代式‘聚焦’机制逐步揭示被隐藏内容。该聚焦是一种渐进式引导策略,使模型通过分步推理逻辑定位图像中的物体,要求层次化注意力转移和动态调整先验认知。本文提出基于策略优化算法的视觉聚焦强化框架,鼓励多模态模型在回答前进行更深入的思考,实现对人类伪装感知系统的对齐甚至超越。大量实验表明,该方法成功催生了多轮推理标记和动态调整检测框的聚焦现象。在伪装物体分类与检测任务中,性能显著优于监督微调(SFT)基线。
原文摘要 · Abstract (English)
Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish concealed objects, demonstrating an inability to emulate human cognitive processes which effectively utilize foreground-background similarity principles for visual analysis. To analyze this hidden human-model visual thinking discrepancy, we build a visual system that mimicks human visual camouflaged perception to progressively and iteratively `refocus' visual concealed content. The refocus is a progressive guidance mechanism enabling models to logically localize objects in visual images through stepwise reasoning. The localization process of concealed objects requires hierarchical attention shifting with dynamic adjustment and refinement of prior cognitive knowledge. In this paper, we propose a visual refocus reinforcement framework via the policy optimization algorithm to encourage multi-modal models to think and refocus more before answering, and achieve excellent reasoning abilities to align and even surpass human camouflaged perception systems. Our extensive experiments on camouflaged perception successfully demonstrate the emergence of refocus visual phenomena, characterized by multiple reasoning tokens and dynamic adjustment of the detection box. Besides, experimental results on both camouflaged object classification and detection tasks exhibit significantly superior performance compared to Supervised Fine-Tuning (SFT) baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。