arXiv:2603.25089cs.CV2026-03中稿 · ICLR

构建真实学术造假场景下的多任务评测基准,检验大模型识破图文欺诈的能力。

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

  • 基于真实撤稿案例与合成数据,设计涵盖7类场景的4000+问题
  • 覆盖5种造假类型和16种细粒度操作,多数样本含多重篡改
  • 从5个核心能力维度评估模型,揭示其优劣势

我们提出THEMIS,一个全新的多任务基准,用于全面评估多模态大语言模型(MLLMs)在真实学术场景中识别视觉欺诈的推理能力。相比现有基准,THEMIS实现三大突破:(1) 真实场景与复杂性:包含超过4000个问题,覆盖7种真实撤稿案例衍生的场景及精心构建的多模态合成数据,其中60.47%为复杂纹理图像,有效弥合了现有基准与真实学术欺诈复杂性之间的差距;(2) 造假类型多样性和细粒度:系统涵盖5种高难度造假类型,引入16种细粒度篡改操作,平均每样本涉及多重叠加操作,对模型的视觉欺诈推理能力提出极高要求;(3) 多维能力评估:建立造假类型到5项核心视觉欺诈推理能力的映射,实现对不同模型在各能力维度上的优劣分析。16个主流MLLM的实验表明,即使表现最佳的GPT-5整体性能也仅达56.15%,证明该基准具有严峻挑战性。我们期望THEMIS能推动MLLM在复杂真实欺诈推理任务中的发展。

原文摘要 · Abstract (English)

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three major advances. (1) Real-World Scenarios and Complexity: Our benchmark comprises over 4,000 questions spanning seven scenarios, derived from authentic retracted-paper cases and carefully curated multimodal synthetic data. With 60.47% complex-texture images, THEMIS bridges the critical gap between existing benchmarks and the complexity of real-world academic fraud. (2) Fraud-Type Diversity and Granularity: THEMIS systematically covers five challenging fraud types and introduces 16 fine-grained manipulation operations. On average, each sample undergoes multiple stacked manipulation operations, with the diversity and difficulty of these manipulations demanding a high level of visual fraud reasoning from the models. (3) Multi-Dimensional Capability Evaluation: We establish a mapping from fraud types to five core visual fraud reasoning capabilities, thereby enabling an evaluation that reveals the distinct strengths and specific weaknesses of different models across these core capabilities. Experiments on 16 leading MLLMs show that even the best-performing model, GPT-5, achieves an overall performance of only 56.15%, demonstrating that our benchmark presents a stringent test. We expect THEMIS to advance the development of MLLMs for complex, real-world fraud reasoning tasks.

多模态模型学术造假评测基准视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。