arXiv:2604.28177cs.CVcs.CY2026-04ACL被引 1

构建首个覆盖多学科的学术图像伪造检测基准,揭示现有方法仍严重滞后。

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

论文配图:AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
图 1 · 摘自论文原文
  • 设计7大领域39类细粒度伪造图像,模拟真实科研场景中的复杂伪造
  • 25种生成模型中11个使检测准确率低于50%,表明反伪造能力普遍不足
  • 首次联合评估检测、推理与定位能力,适合研究学术图像安全的研究者

我们提出AEGIS,一个面向人工智能生成学术图像的全面性评测基准。相比现有基准,AEGIS具有三大创新:(1)领域特异性复杂度:涵盖7个学术类别和39个细粒度子类型,揭示内在的取证难度,即使在GPT-5.1下整体性能也仅达48.80%,专家模型定位准确率(IoU)最高为30.09%;(2)多样化的伪造模拟:在25个生成模型中模拟四种主流学术伪造策略,其中11个模型导致平均取证准确率低于50%,显示取证技术落后于生成进展;(3)多维度取证评估:联合评估检测、推理与定位能力,揭示模型家族间互补优势,多模态大语言模型(MLLMs)在文本伪影识别上达到84.74%准确率,专家检测器在二分类真实性检测中最高达79.54%。通过评估25个主流MLLMs、9个专家模型及1个统一多模态模型,AEGIS成为诊断学术图像取证根本局限性的测试平台。

原文摘要 · Abstract (English)

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) Domain-Specific Complexity: covering seven academic categories with 39 fine-grained subtypes, exposing intrinsic forensic difficulty, where even GPT-5.1 reaches 48.80% overall performance and expert models achieve only limited localization accuracy (IoU 30.09%); (2) Diverse Forgery Simulations: modeling four prevalent academic forgery strategies across 25 generative models, with 11 yielding average forensic accuracy below 50%, showing that forensics lag behind generative advances; and (3) Multi-Dimensional Forensic Evaluation: jointly assessing detection, reasoning, and localization, revealing complementary strengths between model families, with multimodal large language models (MLLMs) at 84.74% accuracy in textual artifact recognition and expert detectors peaking at 79.54% accuracy in binary authenticity detection. By evaluating 25 leading MLLMs, nine expert models, and one unified multimodal understanding and generation model, AEGIS serves as a diagnostic testbed exposing fundamental limitations in academic image forensics.

图像取证AI伪造学术安全多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。