构建首个全面评估大模型伪造检测能力的基准套件
Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models
- 设计覆盖112种伪造类型、6万余道多选题的评测体系
- 测试22个开源及3个闭源大模型,揭示现有能力局限
- 适合关注AI安全、内容可信度研究者使用
AIGC技术快速发展催生了海量虚假媒体,严重威胁社会安全与公共秩序。为应对这一挑战,学界尝试利用大视觉语言模型(LVLMs)开发鲁棒的伪造检测工具。然而,当前缺乏系统性评估框架来全面检验LVLM在伪造识别上的综合能力。为此,本文提出Forensics-Bench,一个涵盖5个维度(伪造语义、模态、任务、类型、模型)、112种伪造类型、共63,292道多选题的综合性评测基准,要求模型具备识别、定位与推理能力。我们对22个开源及3个闭源模型(GPT-4o、Gemini 1.5 Pro、Claude 3.5 Sonnet)进行了全面评估,揭示了当前模型在复杂伪造检测任务中的显著挑战。该基准将持续更新,网址:https://Forensics-Bench.github.io/。
原文摘要 · Abstract (English)
Recently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies have proposed to exploit Large Vision Language Models (LVLMs) to design robust forgery detectors due to their impressive performance on a wide range of multimodal tasks. However, it still lacks a comprehensive benchmark designed to comprehensively assess LVLMs' discerning capabilities on forgery media. To fill this gap, we present Forensics-Bench, a new forgery detection evaluation benchmark suite to assess LVLMs across massive forgery detection tasks, requiring comprehensive recognition, location and reasoning capabilities on diverse forgeries. Forensics-Bench comprises 63,292 meticulously curated multi-choice visual questions, covering 112 unique forgery detection types from 5 perspectives: forgery semantics, forgery modalities, forgery tasks, forgery types and forgery models. We conduct thorough evaluations on 22 open-sourced LVLMs and 3 proprietary models GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet, highlighting the significant challenges of comprehensive forgery detection posed by Forensics-Bench. We anticipate that Forensics-Bench will motivate the community to advance the frontier of LVLMs, striving for all-around forgery detectors in the era of AIGC. The deliverables will be updated at https://Forensics-Bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。