构建首个科学图表溯源检测基准,精准识别抄袭证据与修改方式。
SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

- 按内容复用与形式转换分层设计分类体系
- 含2582对正样本和2541对负样本的混合数据集
- 适合学术诚信检测与多模态模型评估者使用
科学图表常承载研究的核心视觉证据,但图表抄袭尚未成为有基准的多模态评估问题。本文提出SciFigPlag-Bench,一个面向学术文档中科学图表的溯源感知推理基准。不同于通用图像相似性或图像取证基准,该基准评估可疑图表是否复用了特定源图的内容、如何进行变换,以及复用证据出现的位置。我们引入分层分类体系,将复用内容与变换方式分离,涵盖材料保留型复用(如整图、子图复用)与抽象内容复用(如数据重表达、结构重绘)。基于此分类体系,构建包含2,582个正样本对和2,541个负样本对的混合基准,融合真实案例、分类引导的合成样例及视觉相似的负样本。基准支持四项诊断任务:成对检测、来源归因、层级复用类型分类与复用对应定位。对多种视觉-语言模型的实验建立初步基线,揭示细粒度溯源推理、复用类型理解与空间证据定位仍存在持续挑战。
原文摘要 · Abstract (English)
Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal evaluation problem. We present SciFigPlag-Bench, a benchmark for provenance-aware reasoning over scientific figures in scholarly documents. Unlike general image-similarity or image-forensics benchmarks, SciFigPlag-Bench evaluates whether a suspicious figure reuses evidence from a specific source figure, how the reused content has been transformed, and where the reused evidence appears. We introduce a factorized taxonomy that separates what is reused from how it is transformed, covering material-preserving reuse, such as full-figure and subfigure reuse, as well as abstract-content reuse, such as data re-expression and structural redraw. Guided by this taxonomy, we construct a hybrid benchmark with 2,582 positive pairs and 2,541 negative pairs, combining documented real-world cases, taxonomy-guided synthetic examples, and visually similar negatives. The benchmark supports four diagnostic tasks: pairwise detection, source attribution, hierarchical reuse-type classification, and reuse correspondence localization. Experiments with diverse vision-language models establish initial baselines and reveal persistent challenges in fine-grained provenance reasoning, reuse-type understanding, and spatial evidence grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。