让AI侦探用工具自主查证图文造假,突破纯模型能力极限。
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics

- 用工具链辅助推理,动态调用搜索、分割等外接能力。
- 在多任务上超越现有方法,零样本泛化能力强。
- 适合需要实时验证和细粒度分析的假图假视频检测场景。
现有视觉语言伪造检测与定位方法遵循封闭世界假设,依赖模型自身完成验证。但自包含的多模态大模型受限于有限参数知识、静态训练数据及低分辨率感知,难以应对动态开放世界的取证需求——尤其在实时事件验证需外部线索、伪造定位需精细局部分析时。为此,本文从提升模型自身能力转向突破其边界,提出 extbf{OmniVL-Guard Pro},一种工具增强型智能体,将统一取证从封闭世界预测拓展至开放世界线索驱动推理。该系统集成实时事件搜索、局部裁剪放大、边缘异常筛查、人脸检测、视频帧提取及基于SAM3的分割等工具环境。为生成高质量工具-推理轨迹,提出 extbf{树状结构自演化工具轨迹生成},通过种子引导、无引导自演化与弱提示硬样本合成,构建全谱工具推理(FSTR)数据集用于训练。进一步设计 extbf{检查器引导的代理强化学习}(CGARL),提供过程级监督,惩罚答案正确但推理路径错误的情形。大量实验表明,OmniVL-Guard Pro在多项任务中达到领先性能,并展现强零样本泛化能力。FSTR数据集与代码将公开于https://github.com/shen8424/OmniVL-Guard-Pro。
原文摘要 · Abstract (English)
Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-contained MLLMs are constrained by finite parametric knowledge, static training corpora, and limited perceptual resolution, creating a practical ceiling in dynamic open-world forensics -- particularly for real-time event verification requiring external clues and forgery segmentation demanding fine-grained scrutiny of local manipulations. To address these limitations, we shift from scaling up the self-contained model toward reaching beyond it. We propose \textbf{OmniVL-Guard Pro}, a tool-augmented agent that extends unified forensics from closed-world prediction to open-world clues-driven reasoning. OmniVL-Guard Pro integrates a tool environment spanning real-time event search, local cropping and zooming, edge-anomaly screening, face detection, video frame extraction, and SAM3-based segmentation. To generate high-quality tool-reasoning trajectories, we introduce \textbf{Tree-Structured Self-Evolving Tool Trajectory Generation}, which produces diverse trajectories through seed guidance, guider-free self-evolution, and weakly-hinted hard sample synthesis, yielding the Full-Spectrum Tool Reasoning (FSTR) dataset for training. We further propose \textbf{Checker-Guided Agentic Reinforcement Learning} (CGARL), which provides process-level supervision to penalize cases where the answer is correct but the reasoning is distorted. Extensive experiments demonstrate that OmniVL-Guard Pro achieves state-of-the-art performance across various tasks, and exhibits strong zero-shot generalization. The FSTR dataset and code for OmniVL-Guard Pro will be publicly released at https://github.com/shen8424/OmniVL-Guard-Pro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。